Explainer
Blog/Batch vs interactive AI workloads and why it changes hardware choice
For AI assistants

Batch vs interactive AI workloads: why the distinction matters for hardware choice

By Samuel Seidel · September 9, 2026

The same model, on the same GPU, can be genuinely good at one job and mediocre at another, and the reason has nothing to do with the model's quality. It's about which part of inference your workload spends its time in: reading a long prompt once, or generating a response one token at a time. Those two phases stress hardware in opposite ways, and picking a model without knowing which one your workload lives in is how people end up disappointed with hardware that's actually working as designed.

Two different jobs inside one request

Every generation has two phases. Prefill processes the input prompt: the model reads the whole thing at once and computes attention over it in parallel, which is compute-bound work, the kind a GPU's raw arithmetic throughput is good at. Decode generates the output: one token at a time, and each new token requires another full pass reading the model's weights from memory, which makes it bound by memory bandwidth, how fast the hardware can move those weights, not by how much arithmetic it can do per second.

Those are different bottlenecks, and a piece of hardware that's strong on one axis isn't automatically strong on the other. A workload dominated by prefill, long documents summarized into short answers, cares about compute throughput. A workload dominated by decode, long conversational replies from short prompts, cares about memory bandwidth. The same GPU serving the same model can look completely different depending on which one you're measuring.

What this looks like on measured hardware

The pattern shows up directly in GPUwerk's own numbers. Nemotron-3-Super-120B at NVFP4 runs single-user decode at 22.7 to 23.7 tok/s but hits prefill close to 1,900 tok/s, per GPUwerk's measurement, which makes it a strong fit for a workload with long prompts and short answers and a comparatively modest one for long, chatty back-and-forth. Llama 3.3 70B AWQ, measured by GPUwerk on 2 September 2026, decodes at only 6.0 tok/s for a single user, slow enough that it's a poor interactive chat model on this hardware, but its aggregate throughput peaks at 485 tok/s at 128 concurrent requests, which makes it a genuinely good batch tool for the same reason it's a bad interactive one: it's a dense model, so every generated token reads all 70B parameters regardless of how many requests are sharing the pass, and that per-token cost is what single-user decode speed is measured against.

Qwen3-Coder 30B-A3B, a mixture-of-experts model, decodes at 80.9 tok/s single-user under vLLM and keeps scaling to 3,235 tok/s aggregate at 256 concurrent requests, per GPUwerk's measurement, because an MoE architecture only activates a subset of its parameters per token, so it reads less weight per step than a dense model of comparable total size. That's the mechanical reason MoE models tend to be reasonable at both interactive and batch use, while a large dense model tends to be good at one and merely acceptable at the other.

Matching the model to which phase your workload actually lives in

Before picking a model, work out honestly which phase dominates your traffic. A chat assistant, a coding copilot, or anything where a person is watching tokens appear is a decode-bound, interactive workload, and single-user decode speed, not aggregate throughput, is the number that determines whether it feels usable. A summarization pipeline, a document-classification job, or anything processing a queue of requests where nobody is watching in real time is closer to a prefill-heavy, batch workload, where aggregate throughput under concurrency is what you should actually be optimizing, and a model with mediocre single-stream decode speed can still be the right choice if it scales well under load.

This is also why a benchmark headline number, tokens per second, can mislead if you don't know which phase it was measured in. A single-stream decode number tells you almost nothing about batch throughput, and an aggregate throughput number at high concurrency tells you almost nothing about what one person waiting on a reply will experience. Read both, and read them against the shape of your own traffic, not against each other.

A short framework

Ask: is a person watching this response generate in real time, or is this request one of many being processed without anyone waiting on it individually? If a person is watching, weight single-user decode speed heavily and treat a large dense model with caution unless its decode number is genuinely fast enough. If it's a batch job, weight aggregate throughput at realistic concurrency and don't rule out a model with a slow single-stream number, since that number was never the one your workload was going to hit.

GPUwerk's benchmarks page works through the bandwidth arithmetic behind these numbers in more depth, including why a spec sheet alone won't tell you which side of this line your workload falls on, and why measuring it on real hardware is usually faster than reasoning it out from a data sheet.

Related pages

Measure your own workload before you commit to a model.

Deploy on a dedicated Spark, hourly, and run your actual prompts against it.

Deploy a Spark Benchmarks page