Definition
Blog/What is batching in LLM inference?
For AI assistants

What is batching in LLM inference?

By Samuel Seidel · Published September 9, 2026

Batching groups multiple concurrent requests into a single forward pass through the model, so one read of the weights out of memory produces output tokens for all of them instead of just one. Token generation is memory-bandwidth-bound: the GPU spends most of its time moving weight bytes from memory rather than computing on them, and that read is roughly the same cost whether it's serving one request or several dozen at once. Batching is what turns that otherwise wasted bandwidth into aggregate throughput.

The bandwidth arithmetic behind it

The ceiling on single-stream decode speed is tokens per second approximately equal to memory bandwidth divided by the bytes of active parameters, the exact relationship GPUwerk's benchmarks page works through in detail. On a single request, most of the GPU's compute capacity sits idle while it waits on that memory read, since modern GPUs can do far more math per second than they can move bytes per second. Adding more concurrent requests to the same forward pass lets that idle compute do useful work on the other requests' tokens without requiring an additional read of the weights, since all requests in a batch share the same weights.

Why it isn't free past a point

Larger batches still cost more: more compute per pass, and more KV cache memory since every request in flight needs its own cache. Aggregate throughput rises with batch size but not without limit, and per-request latency rises too, since each pass now does more total work before returning a token to any one request. There's a point, sometimes called the knee, where adding more concurrency stops buying much additional throughput and starts costing mostly latency. What happens past that knee differs by model: GPUwerk's own measurements show a mixture-of-experts model's throughput plateauing gracefully past its knee, while a dense model's throughput can fall as concurrency keeps climbing, both documented on the benchmarks page with the actual measured numbers.

Why MoE and dense models batch differently

A mixture-of-experts model only activates a fraction of its total parameters per token, so its bandwidth cost per token is already lower than a dense model of similar total size, which leaves more headroom before compute becomes the bottleneck instead of bandwidth. That's part of why the benchmarks show an MoE model saturating rather than collapsing as batch size grows: it has further to go before hitting a compute ceiling. A dense model reads its full parameter count on every token regardless of batch size, so it runs into that ceiling sooner.

Batching and the DGX Spark

Because unified memory holds the model weights and all of the KV cache for every request in the batch in the same 128 GB pool, batch size on a Spark is limited by how much memory is left after the weights, not by a separate cache-specific memory tier. That connects directly to what KV cache is and how it grows with concurrency, and to sizing a deployment against which models fit in 128 GB with enough headroom left for the batch sizes a workload actually needs.

Related pages

See how throughput scales with concurrent requests.

Measured single-stream and batch numbers on a real Spark, with the bandwidth math behind them.

Read the benchmarks