DGX Spark vs H200 for LLM inference
We rent Sparks, not H200s, so read this the same way as our H100 comparison: if you're serving a shared endpoint with real concurrent traffic, the H200 is the right machine and the Spark is not a substitute for it. The H200 is the biggest bandwidth gap we cover on this site. The question worth answering before you spend money is whether your job actually needs that much machine.
Side-by-side specs
| NVIDIA DGX Spark | NVIDIA H200 SXM | |
|---|---|---|
| Memory | 128 GB LPDDR5x, unified, coherent | 141 GB HBM3e |
| Memory bandwidth | 273 GB/s | ~4.8 TB/s |
| GPU/SoC power | 140 W GB10 chip TDP | Up to 700 W, configurable |
| Architecture | Blackwell (GB10) | Hopper |
| Software stack | Full CUDA: PyTorch, vLLM, TensorRT-LLM, NeMo, Triton | Same CUDA stack, at datacenter scale, plus NVLink/NVSwitch multi-GPU pooling |
| Form factor | Desktop box, single unit, or two networked over 200 GbE | Datacenter accelerator, sold as a card or in an 8-GPU server |
| Rental price | $0.79/hour, GPUwerk, billed per minute | Not cited here, see note below |
Spec figures for the DGX Spark (128 GB, 273 GB/s, 140 W GB10 TDP) are from NVIDIA's own DGX Spark product page. H200 SXM figures (141 GB HBM3e, roughly 4.8 TB/s, up to 700 W configurable) are from NVIDIA's H200 datasheet page. We did not fetch a current third-party H200 on-demand rental rate for this page. H200 availability and pricing are newer and less standardized across providers than H100's, so rather than cite a number that could already be stale, we're pointing you to check a provider's own pricing page directly. Our own $0.79/hour Spark rate is current as published on our pricing page.
Where the Spark loses
Same three places as against the H100, wider gap.
- Memory bandwidth. The H200's roughly 4.8 TB/s against the Spark's 273 GB/s is close to an 18x gap, wider than the H100's roughly 12x, per the NVIDIA figures above. Decode tracks bandwidth almost linearly, so this gap shows up directly in tokens per second.
- Decode throughput on dense models. Our own measurements, run with
vllm bench serveon vLLM 0.27.1 on one of our fleet's Spark nodes on 2 September 2026 (raw logs linked below), put Llama 3.3 70B AWQ at 6.0 tok/s single-stream on the Spark, under the 10 tok/s line where a chat interface starts to feel usable. An H200's bandwidth puts the same class of dense model well past comfortable chat speed, with headroom the H100 doesn't have. - Batch serving. Our benchmark run shows aggregate throughput on Llama 3.3 70B AWQ peaking at 485 tok/s around concurrency 128 on a Spark, then falling to 300 tok/s at concurrency 256 as the node collapses under load. The H200's larger 141 GB HBM3e pool holds more KV cache than the H100's 80 GB, which pushes that concurrency wall out further still. If your job is serving concurrent requests from many users, this is the number that matters.
Where the Spark is enough
The case for the Spark doesn't change because the comparison card got bigger.
- It fits 128 GB. A model under roughly 100 GB, quantised, sits comfortably in the Spark's unified memory. An H200's extra headroom over an H100 (141 GB vs 80 GB) only matters if your model or batch size needs it.
- Single-user or small-team. Our benchmark of Qwen3-Coder 30B-A3B AWQ, the mixture-of-experts model staged on our fleet, measured 80.9 tok/s single-stream on a Spark, run the same day as the Llama figure above (raw log linked below). That's comfortably faster than reading speed for one person, or a small team taking turns.
- Development and evaluation. Standing up a Spark, testing a fine-tune, or running an agent loop against your own endpoint doesn't need H200-class throughput. It needs CUDA parity with production so the code you write today runs unchanged on bigger hardware later, which the Spark has, H200 included.
- Lower cost of entry. GPUwerk's $0.79/hour, billed per minute, is a fraction of any H200 rate we'd expect a provider to charge, though we're not citing a specific number here (see the pricing note above). For workloads that don't need the bandwidth, that gap is the whole argument.
Our measured Spark numbers
Taken with vllm bench serve against vLLM 0.27.1 on one GB10 node of our own fleet, nothing else on its GPU, on 2 September 2026. Full context and methodology on the benchmarks page. We have not run these same benchmarks on an H200; the H200 figures on this page are NVIDIA's published specs, not our own measurements.
| Concurrency | Qwen3-Coder 30B-A3B AWQ aggregate / per request |
Llama 3.3 70B AWQ aggregate / per request |
|---|---|---|
| 1 | 80.9 tok/s / 80.9 | 6.0 tok/s / 6.0 |
| 8 | 845 tok/s / 35.2 | 125 tok/s / 5.2 |
| 32 | 1,679 tok/s / 18.4 | 298 tok/s / 3.9 |
| 64 | 2,421 tok/s / 13.1 | 466 tok/s / 2.8 |
| 128 | 3,010 tok/s / 8.2 | 485 tok/s / 1.4 |
| 256 | 3,235 tok/s / 4.5 | 300 tok/s / 0.4 |
512 tokens in, 256 out, three requests per concurrency slot, zero failed requests at every level. Raw logs: Qwen3-Coder · Llama 3.3 70B.
The mixture-of-experts model saturates past its knee and keeps most of its aggregate throughput; the dense 70B collapses past its knee, aggregate throughput falls as concurrency keeps rising. An H200 pushes that knee out further than an H100 does, on both model types, which is the practical shape of the batch-serving gap described above.
What we'd tell a friend
- Building or serving a shared production endpoint with real concurrent traffic: H200 (or H100, if that fits your budget and memory needs), ours won't cut it.
- Model under 100 GB, one person or a small team, dev and eval work: Spark, at a fraction of what an H200 costs to rent.
- You don't know yet which side you're on: rent a Spark for an hour, run your actual model and prompts, and read your own numbers against the 10 tok/s usability line before committing to H200 spend.
A first engagement can scope which side of that line your workload lands on before you commit to either rate.