MikroTik CRS804 DDQ - The Perfect NVIDIA GB10 Switch

Notice: Page may contain affiliate links for which we may earn a small commission through services like Amazon Affiliates or Skimlinks.

Patrick

Administrator
Staff member
Dec 21, 2010
12,646
6,064
113
I thought I would keep some notes here. MikroTik sent one that arrived Monday. This morning, we already purchased a second one.

We did a main site post on its announcement. It is also a switch I saw a prototype of in Riga in July. At the time, it was seen as a low-cost campus-to-campus switch. I immediately told MikroTik's co-founder that this will sell like hotcakes if they design it for the NVIDIA GB10.

Under the hood, it is similar to the MikroTik CRS812-8DS-2DQ-2DDQ-RM with a few caveats:
  1. All of the ports are QSFP56-DD. That means we get 4x 400GbE ports
  2. It is a half-width switch, so you can install this in half-width racks, albeit it is pretty deep, or you can install it side-by-side. We ordered the side-by-side kit, but I think it actually came in the box.
Here is a quick interface list:
MikroTik CRS804 DDQ Interfaces List QSFP56-DD as 8x 56G.jpg

It is currently running on a single PSU at 28.3W idle and is decently quiet. Of course, if you put 4x 15W optics in there it is going to get much louder and use more power.

For the NVIDIA GB10, the ports mean you can do a QSFP56-DD to 2x QSFP56 breakout DAC and connect eight 200Gbps devices to one switch. That is almost perfect for our 8x NVIDIA GB10 cluster.

I say almost because there is a fairly notable challenge. There are only two 10Gbase-T management ports in addition to the QSFP56-DD ports. That practically means that you are not going to use this for storage networking as well.

With the 4x GB10 cluster, we have been using the CRS812 DDQ and hanging storage off of that. For the 8x GB10 cluster, things are much more challenging since even feeding these systems via their 10Gbase-T ports would require 8x 10Gbase-T plus 80Gbps of storage bandwidth which means 160Gbps in a 10Gbase-T switch. Of course, if you have 2-6 GB10's then this is all super easy since you end up with extra ports open.

To me, a pair of CRS520's is the other interesting option, but I think we are going down the CRS804 DDQ route.

These are available, but since we did a MikroTik video this week, and the second CRS804 DDQ will not arrive until next week, we are probably going to be a bit later on this video than we want. I thought it would be worth letting folks know this is out and ready for GB10 clusters.
 

Joel

Active Member
Jan 30, 2015
888
228
43
44
I’ve been interested in this switch as well; currently running 2 sparks with a DAC; I share the concern about connecting storage to a larger cluster. From what I’ve read on the NVIDIA developer forums, vllm clusters only really work with powers of 2 (2, 4, or 8 nodes), so storage would be a challenge.

Are the 10G ports not capable of being joined to the QSFP network stack at all?

If not, then the setup might get convoluted rather quickly (second >8 port 10G switch)
 

Patrick

Administrator
Staff member
Dec 21, 2010
12,646
6,064
113
I’ve been interested in this switch as well; currently running 2 sparks with a DAC; I share the concern about connecting storage to a larger cluster. From what I’ve read on the NVIDIA developer forums, vllm clusters only really work with powers of 2 (2, 4, or 8 nodes), so storage would be a challenge.

Are the 10G ports not capable of being joined to the QSFP network stack at all?

If not, then the setup might get convoluted rather quickly (second >8 port 10G switch)
So that is effectively what we are doing with the 8x GB10 cluster. The other trick is that on the 10G switch, you ideally want 80Gbps of downstream ports for the GB10's 10Gbase-T, plus you need 80Gbps for storage.

I think another really interesting topology would be to run the GB10 ports as 100Gbps ports to two different switches.
 

T_Minus

Build. Break. Fix. Repeat
Feb 15, 2015
7,906
2,230
113
awesome!

@Patrick have you compared (and I missed it) 4xGB10 cluster vs the M3 MAX 512?
 

Patrick

Administrator
Staff member
Dec 21, 2010
12,646
6,064
113
awesome!

@Patrick have you compared (and I missed it) 4xGB10 cluster vs the M3 MAX 512?
No, but that is kinda the idea with building the 8x GB10 cluster. We can scale to 1, 2, 4, and 8 nodes. Article/ video #1 on this will be in March. It is working already, but there is a bit more optimization needed.
 
  • Like
Reactions: Fzdog2 and T_Minus

TrashMaster

Active Member
Sep 8, 2024
118
87
28
No, but that is kinda the idea with building the 8x GB10 cluster. We can scale to 1, 2, 4, and 8 nodes. Article/ video #1 on this will be in March. It is working already, but there is a bit more optimization needed.
I might be missing it but have you gone through the eye-watering low level minutia of configuring 100+G switches for RDMA and NCCL traffic between nodes? Whatever flow control or frame adjustments are needed for "lossless" ethernetz, etc.
 

nasbdh9

Active Member
Aug 4, 2019
249
166
43
All flow control parameters are now distributed via LLDP extended DCBX. This is completed once the parameters are set on the switch and the network card firmware/system accepts DCBX, without requiring individual settings on each host.
 

Patrick

Administrator
Staff member
Dec 21, 2010
12,646
6,064
113
I might be missing it but have you gone through the eye-watering low level minutia of configuring 100+G switches for RDMA and NCCL traffic between nodes? Whatever flow control or frame adjustments are needed for "lossless" ethernetz, etc.
You might think this is crazy, but MiniMax-M2.5, Claude Code, and so forth can set RoCE up and then verify everything is working.
 
  • Like
Reactions: T_Minus

T_Minus

Build. Break. Fix. Repeat
Feb 15, 2015
7,906
2,230
113
You might think this is crazy, but MiniMax-M2.5, Claude Code, and so forth can set RoCE up and then verify everything is working.
It was fun watching Opus via local agent management system login to the google business account I gave it, create a new project, get api key, update it's own configuration file, and utilize it.

That was a real "woah" moment as I watched it on the monitor zip through google's menus, options, configuration and then update itself LOL!

2026 is going to be wild!
 

Patrick

Administrator
Staff member
Dec 21, 2010
12,646
6,064
113
It was fun watching Opus via local agent management system login to the google business account I gave it, create a new project, get api key, update it's own configuration file, and utilize it.

That was a real "woah" moment as I watched it on the monitor zip through google's menus, options, configuration and then update itself LOL!

2026 is going to be wild!
Just as an experiment, this afternoon we had Sonnet 4.6 build a cluster with eight of the nine GB10's using a new CRS804 DDQ compared to the Apple Mac Studio M3 Ultra 512GB running MiniMax-M2.5. It was benchmarking gpt-oss-120b and minimax-m2.5 by the time we left the studio.
 

Joel

Active Member
Jan 30, 2015
888
228
43
44
Just as an experiment, this afternoon we had Sonnet 4.6 build a cluster with eight of the nine GB10's using a new CRS804 DDQ compared to the Apple Mac Studio M3 Ultra 512GB running MiniMax-M2.5. It was benchmarking gpt-oss-120b and minimax-m2.5 by the time we left the studio.
What's your experience been with an 8 DGX spark cluster, and more importantly, would you still choose it over other options knowing what you do now? Doing the math, it's roughly a $35K expenditure; that's within shouting distance of 4x RTX Pro 6000s.
 

Patrick

Administrator
Staff member
Dec 21, 2010
12,646
6,064
113
What's your experience been with an 8 DGX spark cluster, and more importantly, would you still choose it over other options knowing what you do now? Doing the math, it's roughly a $35K expenditure; that's within shouting distance of 4x RTX Pro 6000s.
I think we are filming that video tomorrow. Quick thoughts:
  • M3 Ultra 512GB is great. If they have a M5 Ultra 1TB, we will probably get one. There is an advantage to just having a lot of compute in one box because you can even just do like LM Studio and be up and running with an OpenAI API endpoint in like 5 minutes. Actually, I use the M3 Ultra as my main studio desktop, and I can usually do my Chrome work and review videos in Premiere Pro without causing an issue for MiniMax-M2.5 running on it. Where it sucks is when you have big prefill jobs. Pro tip there: use the GUFF not the MLX model
  • 4x RTX Pro 6000 - 384GB of memory is very useful. Smaller models will run super fast. Easy since it is all in one box, but that box will likely be >1.44kW
  • The new DGX AI Station - those are going to cost $100-120K, so they are going to be the big step up
  • 8x cluster - Complexity is high. Scaling is not great for small sparse models. You pretty much need to be using a big model to make it worthwhile. Then the question is, are you going to be happy with Kimi K2.5 at 10-20T/s. That seems slow, which is fair, but it is running
So here is where I am at this point:
  • I have a huge bias right now towards using accurate models. Fast and inaccurate is just a pain because it leads to stalled tasks, errors, and so forth.
  • You can see immediately the difference running like Claude Sonnet 4.6 that is hosted and a local machine from a speed perspective. MiniMax-M2.5 8-bit gets you maybe not all the way there, but not super far off
  • Folks running big models at like 3-bit say they are running big models, but the quality goes way down
  • gpt-oss-120b is still useful, but it is an enormous drop off
  • Qwen3.5 122B A10B is becoming my preferred single-node model
  • I like the 4x RTX Pro for running a lot of "smaller" models very fast
  • I like the 8x GB10 cluster for running big models locally
  • M3 Ultra 512GB is big with lower complexity
  • Fast back-end storage helps lower the GB10 cost a lot because you can use 1TB nodes and save $1000 per. That same storage you can use to have different models work on parts of tasks so you get a lot of leverage out of it. Still, if you are downloading 600GB+ models regularly, storage becomes costly.
One of the other strange things is that we were buying the $2999 1TB models, but then prices went up due to the 128GB and the DRAM pricing. When we started this, it was 8x 3K ($24K) for the nodes, 1x $1.1K for the switch, $600 for a decent 10Gbase-T switch, and maybe $600 for the cables, so the cost was more like $26-26.5K. We actually got a 9th node that we ordered, but it got stuck in shipping so it arrived after we set everything else up (our request to cancel was denied). So the box arrived and the price of it would be $650 more to re-order. The question is whether you send it back or not.

At the same time, you have a pretty awesome training cluster if you are fine-tuning models. You also have 160 Arm cores with 200Gbps of networking in the cluster.
 

TrashMaster

Active Member
Sep 8, 2024
118
87
28
I think we are filming that video tomorrow. Quick thoughts:
  • M3 Ultra 512GB is great. If they have a M5 Ultra 1TB, we will probably get one. There is an advantage to just having a lot of compute in one box because you can even just do like LM Studio and be up and running with an OpenAI API endpoint in like 5 minutes. Actually, I use the M3 Ultra as my main studio desktop, and I can usually do my Chrome work and review videos in Premiere Pro without causing an issue for MiniMax-M2.5 running on it. Where it sucks is when you have big prefill jobs. Pro tip there: use the GUFF not the MLX model
  • 4x RTX Pro 6000 - 384GB of memory is very useful. Smaller models will run super fast. Easy since it is all in one box, but that box will likely be >1.44kW
  • The new DGX AI Station - those are going to cost $100-120K, so they are going to be the big step up
  • 8x cluster - Complexity is high. Scaling is not great for small sparse models. You pretty much need to be using a big model to make it worthwhile. Then the question is, are you going to be happy with Kimi K2.5 at 10-20T/s. That seems slow, which is fair, but it is running
So here is where I am at this point:
  • I have a huge bias right now towards using accurate models. Fast and inaccurate is just a pain because it leads to stalled tasks, errors, and so forth.
  • You can see immediately the difference running like Claude Sonnet 4.6 that is hosted and a local machine from a speed perspective. MiniMax-M2.5 8-bit gets you maybe not all the way there, but not super far off
  • Folks running big models at like 3-bit say they are running big models, but the quality goes way down
  • gpt-oss-120b is still useful, but it is an enormous drop off
  • Qwen3.5 122B A10B is becoming my preferred single-node model
  • I like the 4x RTX Pro for running a lot of "smaller" models very fast
  • I like the 8x GB10 cluster for running big models locally
  • M3 Ultra 512GB is big with lower complexity
  • Fast back-end storage helps lower the GB10 cost a lot because you can use 1TB nodes and save $1000 per. That same storage you can use to have different models work on parts of tasks so you get a lot of leverage out of it. Still, if you are downloading 600GB+ models regularly, storage becomes costly.
One of the other strange things is that we were buying the $2999 1TB models, but then prices went up due to the 128GB and the DRAM pricing. When we started this, it was 8x 3K ($24K) for the nodes, 1x $1.1K for the switch, $600 for a decent 10Gbase-T switch, and maybe $600 for the cables, so the cost was more like $26-26.5K. We actually got a 9th node that we ordered, but it got stuck in shipping so it arrived after we set everything else up (our request to cancel was denied). So the box arrived and the price of it would be $650 more to re-order. The question is whether you send it back or not.

At the same time, you have a pretty awesome training cluster if you are fine-tuning models. You also have 160 Arm cores with 200Gbps of networking in the cluster.
I agree that large models (like GLM 4.7 358B FP8) with high precision (fp8/fp16) are infinitely more capable than the smaller <123b stuff, and currently use 4.7 fp8 on a 4x6k dual turin rig. That was, at the time I built it, a ~40k investment, but I can do significant more work, high quality work, and greater complexity of tasks using it. With VLLM and all the correct NCCL, libraries,tweaks, im getting around 6k T/s PP and 60 T/s gen with around 140k context also quantized at fp8.

It would be very useful to see a real PP/TG test with context and full (vllm?) config/tweaks on an 8x cluster of smaller nodes, with instrumentation in place to find the bottlenecks.
image.png

My opinion, right wrong or otherwise, is that in general - most reviews of hardware for 'AI' stuff in the wild today are extremely skewed in favor of catchy clickbait performance metrics of tiny models with very limited real world capabilities. Its neat bob-the-reviewer can get a zillion t/s gen (no pp # presented) on llama3 8b or gemma 3b (unknown quant or inference engine setup), but that's not going to translate to measurable value out the back end. Deeper evaluations to identify understand and optimize the bottlenecks of the platforms are not happening, or not present in the articles/videos/blogs/etc.

My heartfelt request (despite how it might not be the most commercially successful endeavor for professional reviewers) is that the exact testing scenario be well documented and presented, and consistently executed between hardware products for comparison. Platform specific limitations are clearly called out, and understood to be different between prefill and response. Testing of different weight quants is important because based on the platform fp8/16 may either out perform or suck with varying precision. Measure the effective max KV cache available, and explain if the cache is quantized as well.

When I ask about how something is configured, im not asking because I want help setting it up. I have GLM or Kimi or Qwen for that. I am curious about the specific test scenario to understand the results and what they mean past a superficial marketing level. There are communities dedicated to detailed testing and benchmarking on their own hardware to identify and present the most complete view of how to duplicate the results, one such community member from the Blackwell Performance discord/subreddit retains a repo for 6k users: GitHub - voipmonitor/rtx6kpro: RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink worth looking through if you are considering running one of those huge monsters.
 

Joel

Active Member
Jan 30, 2015
888
228
43
44
@TrashMaster Ya it's definitely the wild west about getting accurate information for comparing different hardware. Most youtubers report headline numbers with a prompt of "Tell me a story", don't even look at the output, and that's absolutely useless for real agentic harnesses that feed in 20K tokens of context as part of the system prompt. I've been fairly happy with the 2x DGX sparks using Qwen3.5-122B; that model produces decent code but it's not the best for orchestration IMO.

It seems the audiences are starting to get wise to the real needs of agentic workflows: Moderately high context, high concurrency. There are just so many knobs to turn that it's nearly impossible to get even remotely comparable numbers from different sources. (model/quant/context cache hit rate/concurrency/prompt size, etc)
 

TrashMaster

Active Member
Sep 8, 2024
118
87
28
@Joel funny related bit, I brought this kind of testing up to GN and LTT before in another context and basically got responses like "yeah we test GPUs, but we are not AI guys, so don't really know how to test in depth, meaning we run the cookie cutter test of producing something, even if its not all that useful."

Like TOPs. When somebody talks about TOPs, all I can think of is, wtf is a TOP? How does that equate to my world of actually doing work with the hardware and software at hand? Is that like FLOPs? What kind of FLOP? A 32 bit flop is not a 16 bit flop is not an 8 bit flop is not a 4 bit flop. Marketing numbers throwing out generic FLOPs that turn out to be NVFP4 are like the cake: It's a lie. I'll go right out in the open and say after extensively testing and even producing quants in NVFP4 vs block FP8, that NVFP4 is a lie. We can get into a whole dissertation about how NVFP4 weights require huge amounts of time and resources to re-train properly (because without QAT & QAD they are nowhere near 99.5% similar) and measuring loss is way more complex than just per layer cosine similarity or perplexity.

But I digress.

It's almost like the industry and enthusiast communities need a sample benchmark and platform evaluation to reference. This is what a good AI hardware product review looks like. These are the meat and bones we need to do at a minimum. Not parroting a marketing slide about TOPs or loading lmstudio and downloading llama3 8b with "write me a story" or "write flappy birbz". And to be fair, I am not really criticizing the reviews from STH. They are not bad, they are pretty good by comparison, I just wish they went a little farther and deeper. The pictures on the website of hardware tend to be outstanding by comparison the to trash we get from most reviews.

And to be fair, I understand that this is a huge moving target. Week by week the technology changes, the models change, the benchmarks change, the whole stack will be different in 6 months. But at least for a small point in time, we could compare apples to apples. Pick something midrange, 200-400b at q4_m and fp8, give it a 40k context code file to process and fix a bug. Get real token and context scale and performance, measure the hardware performance to see if its even using the capability of the device its running on.
 
  • Like
Reactions: Mashie and Joel

Patrick

Administrator
Staff member
Dec 21, 2010
12,646
6,064
113
I will tell you, something really tough is just how fast this is moving.

Here is a great example: we have the Beelink ME Pro NAS, and we did the review. I thought it was OK, but it was hard to get excited about. We deleted everything, then had OpenClaw (MiniMax-M2.5 on the M3 Ultra 512GB) just set it up from a base Proxmox VE installation. I saw the video draft today and gave my feedback. It will get published, maybe next week after GTC, but likely the week after. This was pretty mind-blowing in Feb to see it just work that well. By mid-March, will there be a new and better model? We started using MLPerf Client v1.5 just to standardize for client compute platforms, but I feel like it is behind in its suite after a few months.

The GB10 8-node cluster is really interesting because sometimes the scaling just sucks. Other times, it is pretty decent.

One hard part is that I know some folks are doing huge concurrency and so forth to juice tokens/ second. Realistically, I do not think we are anywhere near that in real-world usage on the clusters/ nodes we have. On the other hand, if we do an article/ video and something is getting 30 tokens/ second of useful output, we literally get folks commenting that they are running some like 2-bit quant of some other model on their RTX 3090 faster. So the way to combat that is to run a GB10 or similar device under a scenario that is more like it is running in a giant data center with high loading to get big numbers.

Then again, at some point, I wonder if you have an 8x GB10 cluster: Is that for one person? Is it for a workgroup? Is a single GB10 running Qwen3.5 or gpt-oss-120b something for a workgroup that sits on someone's desk? I know we have many of the AMD Ryzen AI Max+ 395 nodes just running dedicated models, so that they are available and not running on the bigger nodes.

To me, the big value is in being able to get results that are useful. Once you have that, you can scale the compute to be faster. Another way to put it, perhaps, if you spend $30K on a solution that is 1/8th the speed of a $100K, one school of thought is that it is bad since your $/token is worse. I look at it as if that $30K is generating $60K/year of value; then it is worth it, and it might have proved the case for the $100K option as well.

Again, I know that is an odd way to look at it, but I also feel like there are fun and lower-cost setups that are interesting, but at some point, it is about getting a positive return. On bigger setups, I think the goal needs to be getting a setup that directly provides value.
 

T_Minus

Build. Break. Fix. Repeat
Feb 15, 2015
7,906
2,230
113
The AI enthusiast community feels very similar to the SEO Industry ~2005 and the web design industry before that... FULL of charlatans right now.

I would almost go so far as to say 90% or more of users interacting on AI related groups, subs, forums and making YouTube videos know 1% of what they are talking about, it's usually "look what I setup" or "look at this output" both of which may be impressive to non-technical persons, but are completely meaningless with no meat and potatoes of how it was done. The how consists of "type this command, follow the guide guys you can do it!" - I even attempted to share experiences with models and some pointers and was quickly told I was making it all up... or people comparing to 8b models... right... I gave up participating!


I know I have purchased a lot lately to experiment, learn, and figure out the ROI later... from mac mini to 512gb ultra to desktop GPUs... there's a use for them all IMO. From personal, home use (image analysis), to business use (coding, QA, and more).


It feels like it's the 90s again with something new daily, bigger something new weekly, even bigger monthly... it's very fun to be a part of this right now, in my own little way :)
 
Last edited:
  • Like
Reactions: Patrick and Joel

Patrick

Administrator
Staff member
Dec 21, 2010
12,646
6,064
113
My bigger one right now is how many people are still on the AI is only for slop bandwagon. Sure, there is a lot of slop, but also just seeing it turn the corner into something useful on a daily basis makes it hard to interact at that slop-only level
 

TrashMaster

Active Member
Sep 8, 2024
118
87
28
How can we help with performance testing and optimizing the 8 node setup? There is a deep well of knowledge between the forums at STH, L1T, and various subreddits and discords dedicated to productive use of LLMs at scale.

If someone was reviewing an 8x GPU MGX, I might expect batch parallel token processing numbers to be a more valuable metric, but in the enthusiast space I feel like the target is responsive 1 or 2 concurrent request performance information.

You are right that everything is moving too fast, far too fast for a month to pass between recording a bleeding edge video and its release. This dramatically impacts the quality of the videos that can be released without sufficient time to review/edit/revise. Every weekend somebody is vibe coding a new tool/library/innovation. By Monday its in dev. By Wednesday its in production. By the next week, its old hat already on its way out the door lol.

I don't know if the solution is shorter format content only a couple minutes that focuses on the meat and potatoes, or a couple test-environment scenarios explained in depth and simply referenced whenever the next scale-out Nx $thing cluster comes out. The software changes alone result in tremendous performance fluctuations that invalidate results within a couple months and going back to re-test a whole battery of stuff is just not viable when cuda 13.2 moves to cuda 13.3...
 

Joel

Active Member
Jan 30, 2015
888
228
43
44
How can we help with performance testing and optimizing the 8 node setup? There is a deep well of knowledge between the forums at STH, L1T, and various subreddits and discords dedicated to productive use of LLMs at scale.

If someone was reviewing an 8x GPU MGX, I might expect batch parallel token processing numbers to be a more valuable metric, but in the enthusiast space I feel like the target is responsive 1 or 2 concurrent request performance information.
Seems to be the best source of info on DGX Spark is the NVIDIA developer forums.

Regarding concurrency, the latest thing is agents orchestrating subagents for tasks that are able to be done in parallel; I can certainly envision high concurrency being the #1 optimization priority for most users in 6-12 months.