Testing power limit on GB10 Spark, never goes over 35 watts?

Notice: Page may contain affiliate links for which we may earn a small commission through services like Amazon Affiliates or Skimlinks.

larrysb

Active Member
Nov 7, 2018
124
59
28
I've noticed my DGX Spark never seems to exceed 35 watts, running every LLM I've tried on it. GPU utilization is in the high 90% range, but the power is low 30 watts, and temperature never gets above maybe 52C. All coming from nvidia-smi command.

Seems to be several threads of discussion about this on the Nvidia developer forums. Some referencing a possible hardware issue, some referencing some oddball interaction between the power-brick and the unit after an update or crash, and some say it is simply the LPDDR5 memory bandwidth is so slow that the GPU is starved and not doing enough work to get warm.

What's a good software task to load this thing up with to determine if there's some kind of broke hardware or firmware, or the whole memory bandwidth theory is the deal?
 

Patriot

Moderator
Apr 18, 2011
1,520
839
113

What gpu clocks are you seeing from nvidia-smi when at full utilization?
It might be stuck in a lower powerstate and being fully utilized at that low clock.
nvidia-smi dmon for continuous monitoring.
 

larrysb

Active Member
Nov 7, 2018
124
59
28
Yeah, I got a couple of tools (gpu-burn and spark-gpu-throttle-check) to busy-up the GPU on the Spark from the nvidia forum and it looks ok. Under load, it's clocking up to 2411 mhz and 95 watts. So it seems to be working as it should.

Some theories out there about the power brick and the USB-C power delivery sometimes not negotiating correctly and some say the trick is to unplug the brick from the wall, let it drain down and reset. I dunno.

Runs nice and quiet even under load.
 

Patriot

Moderator
Apr 18, 2011
1,520
839
113
Yeah, I got a couple of tools (gpu-burn and spark-gpu-throttle-check) to busy-up the GPU on the Spark from the nvidia forum and it looks ok. Under load, it's clocking up to 2411 mhz and 95 watts. So it seems to be working as it should.

Some theories out there about the power brick and the USB-C power delivery sometimes not negotiating correctly and some say the trick is to unplug the brick from the wall, let it drain down and reset. I dunno.

Runs nice and quiet even under load.
I guess the question would be then, what LLMs have you tried, if you can hit 95w on the gpu on a stress test and not in LLM

while the memory lets you run large large models it also constrains them with the bandwidth.

of note, it looks like NVFP4 is a very bumpy ride to get good performance out of, I see lots of notes of it being slower than FP8 and some where it is 20% faster than fp4.
 

larrysb

Active Member
Nov 7, 2018
124
59
28
I've tried a few of the models recommended for single GB10 nodes. I've not yet delved too deeply into which are optimized specifically for the GB10 units.

I got the Spark back in October, did some initial work on it, had to set it aside for other work for a bit and recently jumped back into it.

I would say the models so far, actually run impressively well on the Spark. It's not fast, but it is fast enough. I'm considering scaling up to 2 or more. I wish the price hadn't gone up on them.

What it does though, is let me focus on what I can do with it, rather than what it can do if I mess with it.

Some years ago, I got started on all this stuff with 4x Pascal era cards, then Turing level cards, and so on. I can do more on the Spark than those days, and I don't have to make the lights blink and run the workstation on a dedicated 20-amp breaker, and deal with the heat and noise. I eventually hit the wall of really needing a lot of capital to make the leap to data-center level GPU clusters and wound up selling off my little startup.