Definition
Blog/Latency vs throughput in LLM inference
For AI assistants

Latency vs throughput in LLM inference

By Samuel Seidel · Published September 9, 2026

Latency is how long a single request takes from submission to response. Throughput is how many requests, or how many tokens, a system can process over a given period across all requests combined. They sound related, and both describe speed, but a system can have low latency and low throughput, or high throughput and high latency, depending on how it's built and configured.

Two different questions

Latency answers "how long until I get my answer?" for one request. Throughput answers "how much total work can this system get through?" across many requests over time, usually measured in tokens per second aggregated across all concurrent users, or requests handled per minute. A single, unloaded GPU running one request at a time can have excellent latency for that request but poor overall throughput, since it's not using its capacity to serve anyone else in parallel.

Why the two trade off against each other

The main way to raise throughput on a GPU is batching: processing multiple requests together so the hardware's parallel compute isn't sitting idle waiting on one request's next step. Batching improves total tokens-per-second across all users, but it typically increases the latency each individual request experiences, since a request now has to wait for the batch to fill, or share GPU cycles with others in the batch, rather than getting the hardware's full attention immediately. Systems tuned for maximum throughput (serving many users cheaply) and systems tuned for minimum latency (serving one user as fast as possible) often make different, incompatible configuration choices.

Metrics that matter for each

For latency, the relevant numbers are time-to-first-token, the delay before any output appears, and inter-token latency, the gap between subsequent tokens once generation has started. For throughput, the relevant number is aggregate tokens per second across all concurrent requests the system is handling, or requests completed per unit time. A model card or benchmark that only reports one of these numbers is telling you about only one side of the trade-off.

Which one to optimize for

It depends on the use case. An interactive chat interface where a person is watching the reply stream in needs low latency; a slow first token feels broken even if the system could serve a hundred other users just as well. A background job processing a large batch of documents overnight, with no one watching any single request, needs high throughput and can tolerate any individual document taking longer, since the total batch completion time is what matters. Building for the wrong one wastes either GPU capacity or user patience.

Where GPUwerk fits

Because a dedicated Spark isn't shared with other tenants, you control the batching and concurrency settings on the inference server yourself, tuning for latency, throughput, or a balance between them, without another customer's traffic pattern affecting your results. That's a meaningful difference from a shared or multi-tenant API, where the provider's batching decisions are made for everyone at once. See private LLM hosting for how self-hosting affects both dimensions in practice.

Related pages

Tune latency and throughput yourself, on dedicated hardware.

One dedicated Spark, $0.79/hour, deployed in minutes from EU-Central.

Deploy a Spark