DGX Spark vs H100 for LLM inference
We rent Sparks, not H100s, so read the next sentence as us arguing against our own product where it matters: if you're serving a shared endpoint with real concurrent traffic, the H100 is the right machine and the Spark is not a substitute for it. The Spark is enough for a narrower job, and this page is about finding out which side of that line you're on before you spend money.
Side-by-side specs
| NVIDIA DGX Spark | NVIDIA H100 SXM | |
|---|---|---|
| Memory | 128 GB LPDDR5x, unified, coherent | 80 GB HBM3 |
| Memory bandwidth | 273 GB/s | 3.35 TB/s |
| GPU/SoC power | 140 W GB10 chip TDP | Up to 700 W, configurable |
| Software stack | Full CUDA: PyTorch, vLLM, TensorRT-LLM, NeMo, Triton | Same CUDA stack, at datacenter scale, plus NVLink/NVSwitch multi-GPU pooling |
| Form factor | Desktop box, single unit, or two networked over 200 GbE | Datacenter accelerator, sold as a card or in an 8-GPU server |
| Rental price | $0.79/hour, GPUwerk, billed per minute | From $1.99/hour (PCIe) to $2.69/hour (SXM), Runpod on-demand, fetched Sept 7, 2026 |
Spec figures for the DGX Spark (128 GB, 273 GB/s, 140 W GB10 TDP) are from NVIDIA's own DGX Spark product page, fetched September 7, 2026. H100 SXM figures (80 GB HBM3, 3.35 TB/s, up to 700 W configurable) are from NVIDIA's H100 datasheet page, fetched the same day. H100 rental prices are from Runpod's pricing page, fetched September 7, 2026: $1.99/hour for H100 PCIe and $2.69/hour for H100 SXM on-demand, both before any reserved or spot discount.
Where the Spark loses
Three places, and they're related.
- Memory bandwidth. The H100's 3.35 TB/s against the Spark's 273 GB/s is a roughly 12x gap, per the NVIDIA figures above. Decode is a bandwidth-bound problem: every generated token means reading the active weights out of memory once, so single-stream generation speed tracks bandwidth almost linearly. Nothing in software closes a 12x hardware gap.
- Decode throughput on dense models. Our own measurements, run with
vllm bench serveon vLLM 0.27.1 on one of our fleet's Spark nodes on 2 September 2026 (raw logs linked below), put Llama 3.3 70B AWQ at 6.0 tok/s single-stream on the Spark, under the 10 tok/s line where a chat interface starts to feel usable. An H100's bandwidth puts the same class of dense model well into comfortable chat-speed territory. - Batch serving. The same benchmark run shows why concurrency doesn't rescue a dense model on a Spark: aggregate throughput on Llama 3.3 70B AWQ peaks at 485 tok/s around concurrency 128, then falls to 300 tok/s at concurrency 256 as the node collapses under load. An H100, with more bandwidth and a larger memory pool to hold KV cache, sustains far higher aggregate throughput before hitting that kind of wall. If your job is serving concurrent requests from many users, this is the number that matters, not any single-stream figure.
Where the Spark is enough
The same numbers read differently for a narrower job.
- It fits 128 GB. A model under roughly 100 GB, quantised, sits comfortably in the Spark's unified memory with room to spare, and you don't need an 80 GB H100's tighter budget or multi-GPU sharding to load it.
- Single-user or small-team. Our benchmark of Qwen3-Coder 30B-A3B AWQ, the mixture-of-experts model staged on our fleet, measured 80.9 tok/s single-stream on a Spark, run the same day as the Llama figure above (raw log linked below). That's comfortably faster than reading speed for one person, or a small team taking turns.
- Development and evaluation. Standing up a Spark, testing a fine-tune, checking whether a quantisation format holds output quality, or running an agent loop against your own endpoint doesn't need H100-class throughput. It needs CUDA parity with production so the code you write today runs unchanged on bigger hardware later, which the Spark has.
- Per-minute billing at a lower rate. GPUwerk's $0.79/hour, billed per minute, against Runpod's $1.99/hour to $2.69/hour on-demand H100 rate (fetched September 7, 2026) is roughly a third of the cost for workloads that don't need the extra bandwidth.
Our measured Spark numbers
Taken with vllm bench serve against vLLM 0.27.1 on one GB10 node of our own fleet, nothing else on its GPU, on 2 September 2026. Full context and methodology on the benchmarks page.
| Concurrency | Qwen3-Coder 30B-A3B AWQ aggregate / per request |
Llama 3.3 70B AWQ aggregate / per request |
|---|---|---|
| 1 | 80.9 tok/s / 80.9 | 6.0 tok/s / 6.0 |
| 8 | 845 tok/s / 35.2 | 125 tok/s / 5.2 |
| 32 | 1,679 tok/s / 18.4 | 298 tok/s / 3.9 |
| 64 | 2,421 tok/s / 13.1 | 466 tok/s / 2.8 |
| 128 | 3,010 tok/s / 8.2 | 485 tok/s / 1.4 |
| 256 | 3,235 tok/s / 4.5 | 300 tok/s / 0.4 |
512 tokens in, 256 out, three requests per concurrency slot, zero failed requests at every level. Raw logs: Qwen3-Coder · Llama 3.3 70B.
The mixture-of-experts model saturates past its knee and keeps most of its aggregate throughput; the dense 70B collapses past its knee, aggregate throughput falls as concurrency keeps rising. An H100 pushes that knee much further out on both model types, which is the practical shape of the batch-serving gap described above.
What we'd tell a friend
- Building or serving a shared production endpoint with real concurrent traffic: H100, ours won't cut it.
- Model under 100 GB, one person or a small team, dev and eval work: Spark, at a third of the hourly cost.
- You don't know yet which side you're on: rent a Spark for an hour, run your actual model and prompts, and read your own numbers against the 10 tok/s usability line before committing to H100 spend.
A first engagement can scope which side of that line your workload lands on before you commit to either rate.