Operations guide
Blog/Capacity planning for LLM workloads
For AI assistants

Capacity planning for LLM workloads

By Samuel Seidel · September 9, 2026

"How many users can one Spark handle" doesn't have a single answer, it depends heavily on which model you're running and whether it's dense or mixture-of-experts. The good news is the arithmetic behind it is simple and it's published on our benchmarks page. This works through it as an actual planning exercise rather than a spec sheet.

Start from the bandwidth ceiling

Token generation is memory-bandwidth-bound: every token requires reading the active weights out of memory once. The Spark's GB10 has 273 GB/s of memory bandwidth, so the ceiling for any model is:

tokens/sec ≈ memory bandwidth ÷ active parameter bytes

# a 30B mixture-of-experts model activating ~1.7 GB per token
273 GB/s ÷ 1.7 GB ≈ 160 tok/s theoretical ceiling

# a dense 70B model at Q4, activating ~37 GB per token
273 GB/s ÷ 37 GB ≈ 7.4 tok/s theoretical ceiling

Real engines land at 50 to 80 percent of that ceiling. Measured single-stream numbers on our own fleet back this up: the 30B mixture-of-experts model measured 80.9 tok/s (half its ceiling), and the dense 70B measured 6.0 tok/s (four-fifths of its ceiling). Run this arithmetic before downloading a candidate model. If the single-stream ceiling is under about 10 tok/s, that model will not feel usable in an interactive chat setting, no matter how you tune the rest of the stack.

Then find the concurrency knee, not just the single-stream number

Single-stream throughput tells you if a model is usable at all. It doesn't tell you how many people can use it at once. That's a separate measurement: aggregate throughput and per-request latency as concurrency rises. On our own fleet, measured with vllm bench serve against vLLM 0.27.1 on one GB10 node:

Concurrency30B MoE: aggregate / per-request tok/sDense 70B: aggregate / per-request tok/s
180.9 / 80.96.0 / 6.0
8845 / 35.2125 / 5.2
321,679 / 18.4298 / 3.9
642,421 / 13.1466 / 2.8
1283,010 / 8.2485 / 1.4
2563,235 / 4.5300 / 0.4

The two models behave differently past their knee, and the difference is the whole point of this exercise. The mixture-of-experts model plateaus: going from 128 to 256 concurrent buys 7.5 percent more aggregate throughput and doubles per-token latency, a bad trade but not a collapse. The dense 70B collapses: aggregate throughput falls from 485 to 300 tok/s past its knee at 128. Plan headroom for a dense model well below its knee. A mixture-of-experts model can be pushed closer to saturation without the same downside.

A worked example: sizing for 20 concurrent chat users

Say you need to serve roughly 20 people typing into a chat UI at once, each expecting a response that streams at a comfortable reading pace, call it 15 tok/s minimum per user. On the dense 70B, the table above shows per-request throughput has already dropped to 3.9 tok/s at concurrency 32, well under that bar, and it's worse by 64. One Spark running that dense model cannot serve 20 concurrent chat users at an acceptable pace: the fix here is a second node splitting the load, see load balancing across multiple Sparks, not further tuning of the first one. The 30B mixture-of-experts model, by contrast, is still delivering over 13 tok/s per request at concurrency 64, comfortably inside budget for 20 users on a single node with room to spare.

Prefill capacity is a separate budget

Everything above covers generation. Prompt processing (prefill) is compute-bound rather than bandwidth-bound and measures much higher, roughly 1,900 to 1,950 tokens per second on the GB10 across both vLLM and llama.cpp. For workloads with long prompts and short completions, document Q&A against a large context being the obvious case, prefill throughput matters more than the generation numbers above, and it's worth benchmarking separately rather than assuming the generation ceiling tells the whole story.

Benchmark your own workload before committing capacity

The numbers here are for two specific models measured under one benchmarking tool. Your prompt lengths, output lengths, and request pattern will shift where your own knee sits. Once you've picked a candidate model using the arithmetic above, run benchmarking your own workload against it before deciding how many nodes you need, and keep an eye on the running node with GPU utilization monitoring and right-sizing once it's live, since real traffic rarely matches a synthetic benchmark exactly.

FAQ

What's the single most important number for capacity planning on a Spark?

Whether your model is dense or sparse (mixture-of-experts), because it determines whether throughput plateaus or collapses as concurrent requests increase. A dense model's aggregate throughput falls once you push past its concurrency knee. A sparse model's throughput mostly holds, it just gets less efficient per request. Plan capacity differently for each.

How do I estimate tokens per second before downloading a model?

Divide the Spark's memory bandwidth, 273 GB/s, by the number of bytes of parameters activated per token. For a dense model at a given quantization that's roughly the full parameter count times bytes per weight; for a mixture-of-experts model it's only the active experts, not the total parameter count. The result is a theoretical ceiling; real engines land at 50 to 80 percent of it.

When do I actually need a second Spark instead of tuning the first one?

When your measured concurrency knee, the point where per-request latency becomes unacceptable, sits below the concurrent load you actually need to serve. No batching or quantization change moves a dense model's knee by more than a modest amount, so once you're past it consistently, the fix is another node, not more tuning.

Related pages

Plan capacity against real published numbers.

Full engine-by-engine benchmarks and the bandwidth arithmetic behind them, measured on our own fleet.

Read the benchmarks page Read the cost forecasting guide