Comparison

DGX Spark vs H200 for LLM inference

By Samuel Seidel · Updated September 9, 2026 · 7 min read

We rent Sparks, not H200s, so read this the same way as our H100 comparison: if you're serving a shared endpoint with real concurrent traffic, the H200 is the right machine and the Spark is not a substitute for it. The H200 is the biggest bandwidth gap we cover on this site. The question worth answering before you spend money is whether your job actually needs that much machine.

Side-by-side specs

 NVIDIA DGX SparkNVIDIA H200 SXM
Memory128 GB LPDDR5x, unified, coherent141 GB HBM3e
Memory bandwidth273 GB/s~4.8 TB/s
GPU/SoC power140 W GB10 chip TDPUp to 700 W, configurable
ArchitectureBlackwell (GB10)Hopper
Software stackFull CUDA: PyTorch, vLLM, TensorRT-LLM, NeMo, TritonSame CUDA stack, at datacenter scale, plus NVLink/NVSwitch multi-GPU pooling
Form factorDesktop box, single unit, or two networked over 200 GbEDatacenter accelerator, sold as a card or in an 8-GPU server
Rental price$0.79/hour, GPUwerk, billed per minuteNot cited here, see note below

Spec figures for the DGX Spark (128 GB, 273 GB/s, 140 W GB10 TDP) are from NVIDIA's own DGX Spark product page. H200 SXM figures (141 GB HBM3e, roughly 4.8 TB/s, up to 700 W configurable) are from NVIDIA's H200 datasheet page. We did not fetch a current third-party H200 on-demand rental rate for this page. H200 availability and pricing are newer and less standardized across providers than H100's, so rather than cite a number that could already be stale, we're pointing you to check a provider's own pricing page directly. Our own $0.79/hour Spark rate is current as published on our pricing page.

Where the Spark loses

Same three places as against the H100, wider gap.

Where the Spark is enough

The case for the Spark doesn't change because the comparison card got bigger.

Our measured Spark numbers

Taken with vllm bench serve against vLLM 0.27.1 on one GB10 node of our own fleet, nothing else on its GPU, on 2 September 2026. Full context and methodology on the benchmarks page. We have not run these same benchmarks on an H200; the H200 figures on this page are NVIDIA's published specs, not our own measurements.

Concurrency Qwen3-Coder 30B-A3B AWQ
aggregate / per request
Llama 3.3 70B AWQ
aggregate / per request
1 80.9 tok/s / 80.9 6.0 tok/s / 6.0
8 845 tok/s / 35.2 125 tok/s / 5.2
32 1,679 tok/s / 18.4 298 tok/s / 3.9
64 2,421 tok/s / 13.1 466 tok/s / 2.8
128 3,010 tok/s / 8.2 485 tok/s / 1.4
256 3,235 tok/s / 4.5 300 tok/s / 0.4

512 tokens in, 256 out, three requests per concurrency slot, zero failed requests at every level. Raw logs: Qwen3-Coder · Llama 3.3 70B.

The mixture-of-experts model saturates past its knee and keeps most of its aggregate throughput; the dense 70B collapses past its knee, aggregate throughput falls as concurrency keeps rising. An H200 pushes that knee out further than an H100 does, on both model types, which is the practical shape of the batch-serving gap described above.

What we'd tell a friend

A first engagement can scope which side of that line your workload lands on before you commit to either rate.

First top-up: pay $10, get $20 in credit

Find out which side you're on for $0.79/hour.

Run your real model on a dedicated DGX Spark before you commit to H200 spend, ours or anyone else's.

Deploy a Spark See pricing