Comparison

DGX Spark vs H100 for LLM inference

By Samuel Seidel · Updated September 7, 2026 · 8 min read

We rent Sparks, not H100s, so read the next sentence as us arguing against our own product where it matters: if you're serving a shared endpoint with real concurrent traffic, the H100 is the right machine and the Spark is not a substitute for it. The Spark is enough for a narrower job, and this page is about finding out which side of that line you're on before you spend money.

Side-by-side specs

 NVIDIA DGX SparkNVIDIA H100 SXM
Memory128 GB LPDDR5x, unified, coherent80 GB HBM3
Memory bandwidth273 GB/s3.35 TB/s
GPU/SoC power140 W GB10 chip TDPUp to 700 W, configurable
Software stackFull CUDA: PyTorch, vLLM, TensorRT-LLM, NeMo, TritonSame CUDA stack, at datacenter scale, plus NVLink/NVSwitch multi-GPU pooling
Form factorDesktop box, single unit, or two networked over 200 GbEDatacenter accelerator, sold as a card or in an 8-GPU server
Rental price$0.79/hour, GPUwerk, billed per minuteFrom $1.99/hour (PCIe) to $2.69/hour (SXM), Runpod on-demand, fetched Sept 7, 2026

Spec figures for the DGX Spark (128 GB, 273 GB/s, 140 W GB10 TDP) are from NVIDIA's own DGX Spark product page, fetched September 7, 2026. H100 SXM figures (80 GB HBM3, 3.35 TB/s, up to 700 W configurable) are from NVIDIA's H100 datasheet page, fetched the same day. H100 rental prices are from Runpod's pricing page, fetched September 7, 2026: $1.99/hour for H100 PCIe and $2.69/hour for H100 SXM on-demand, both before any reserved or spot discount.

Where the Spark loses

Three places, and they're related.

Where the Spark is enough

The same numbers read differently for a narrower job.

Our measured Spark numbers

Taken with vllm bench serve against vLLM 0.27.1 on one GB10 node of our own fleet, nothing else on its GPU, on 2 September 2026. Full context and methodology on the benchmarks page.

Concurrency Qwen3-Coder 30B-A3B AWQ
aggregate / per request
Llama 3.3 70B AWQ
aggregate / per request
1 80.9 tok/s / 80.9 6.0 tok/s / 6.0
8 845 tok/s / 35.2 125 tok/s / 5.2
32 1,679 tok/s / 18.4 298 tok/s / 3.9
64 2,421 tok/s / 13.1 466 tok/s / 2.8
128 3,010 tok/s / 8.2 485 tok/s / 1.4
256 3,235 tok/s / 4.5 300 tok/s / 0.4

512 tokens in, 256 out, three requests per concurrency slot, zero failed requests at every level. Raw logs: Qwen3-Coder · Llama 3.3 70B.

The mixture-of-experts model saturates past its knee and keeps most of its aggregate throughput; the dense 70B collapses past its knee, aggregate throughput falls as concurrency keeps rising. An H100 pushes that knee much further out on both model types, which is the practical shape of the batch-serving gap described above.

What we'd tell a friend

A first engagement can scope which side of that line your workload lands on before you commit to either rate.

First top-up: pay $10, get $20 in credit

Find out which side you're on for $0.79/hour.

Run your real model on a dedicated DGX Spark before you commit to H100 spend, ours or anyone else's.

Deploy a Spark See pricing