Load-testing your inference endpoint before production
A single curl request that returns a fast, correct completion tells you the model loaded and the engine works. It tells you almost nothing about what happens when fifty of those requests arrive in the same second. Load testing closes that gap, and for a self-hosted endpoint you're operating yourself, it's the step most likely to get skipped because the single-request test already "worked."
What breaks under concurrency that doesn't break under a single request
KV cache exhaustion is the most common one: memory that comfortably fits one request's context can run out once several concurrent requests are each holding their own cache, a failure mode covered in more depth in troubleshooting OOM errors. Latency also doesn't scale linearly; a request that takes two seconds alone can take considerably longer once it's queued behind others competing for the same GPU, and the relationship between concurrency and per-request latency is worth measuring directly rather than assumed. Finally, anything stateful in front of the engine, a reverse proxy, a rate limiter, an authentication check against a database, can become the actual bottleneck under load even when the inference engine itself has headroom left.
Tools for generating concurrent load
vLLM ships a benchmarking script (`benchmark_serving.py` in its repository) built specifically for OpenAI-compatible completion endpoints, which handles concurrency, request pacing, and latency percentile reporting without you having to build that tooling yourself. For a more general-purpose HTTP load tool, `hey`, `wrk`, or `k6` all work against any HTTP API and are worth using if you want load patterns vLLM's own script doesn't cover, such as testing through your full stack including a reverse proxy or LiteLLM instance rather than hitting the engine directly. Testing through the full stack is the more representative test, since it's what production traffic will actually go through, rather than the bare engine alone.
What to measure beyond pass or fail
Latency at multiple percentiles, not only the average: p50 tells you what a typical user experiences, p99 tells you what your worst-off users experience, and the gap between them tends to widen as concurrency increases, which an average alone hides. Throughput, requests or tokens served per second, at increasing concurrency levels until it plateaus or degrades, to find where the endpoint's actual capacity ceiling is rather than assuming it matches a spec-sheet number. Error rate at each concurrency level, since a system that degrades by slowing down is a very different failure mode from one that degrades by returning errors, and which one your setup does matters for how the calling application should handle it.
Test the request shapes you'll actually serve
A load test built entirely from short, similar prompts will not reveal how the endpoint behaves under a realistic mix of short chat messages and occasional long documents. If your real traffic includes long-context requests, include them in the load test explicitly, per handling long-context requests, since KV cache pressure from a long request under concurrent load is exactly the scenario short-prompt testing misses. Same for streaming versus non-streaming, discussed in streaming vs non-streaming responses: a streaming connection held open for the duration of generation behaves differently under high concurrency than a request-response cycle, and testing only one mode won't tell you about the other.
Find the ceiling, then set limits below it
The point of a load test isn't just confirming the endpoint survives your expected traffic; it's finding where it stops surviving, so you can set concurrency limits comfortably below that ceiling rather than discovering it in production. Once you know the concurrency level where latency degrades unacceptably or errors start appearing, configure LiteLLM or your reverse proxy to cap concurrent requests below that point, covered in our LiteLLM guide, so the system degrades by queueing or rejecting excess requests cleanly instead of falling over for everyone once the real ceiling is hit.
Retest after any change to the stack
A load test result is specific to the model, quantization, engine version, and hardware it was run against. Changing the quantization level, per choosing a quantization level, upgrading the inference engine, or moving to a different node all invalidate a prior load test's numbers, since any of them can shift the actual capacity ceiling up or down. Treat load testing as part of the deployment process for a meaningful stack change, not a one-time exercise done before initial launch and never revisited.
FAQ
How is load testing different from benchmarking a model?
Benchmarking, covered in our guide on benchmarking your own workload, measures how fast a model runs under a controlled, usually single-stream, workload. Load testing measures how the whole endpoint, engine, proxy, and network, behaves under many concurrent requests arriving the way real traffic would, including how it degrades once demand exceeds capacity. A model can benchmark well in isolation and still fall over under concurrent load if the endpoint wasn't tested for it.
What concurrency level should I load-test at?
Start with your actual expected peak concurrent users or requests, then test somewhat above that, since real traffic is rarely perfectly smooth and peaks exceed the average. There's no universal target number; it depends entirely on your application's traffic pattern, which is why load testing with synthetic numbers pulled from an unrelated project's results is not a substitute for testing your own expected load.
Can I load-test against a shared or idle-tier endpoint and trust the results?
No. Results from a shared node reflect whatever else was running on it at the time, not the capacity you'll actually have in production, and results from an idle-tier or best-effort deployment reflect degraded conditions that won't match a dedicated node under normal load. Load-test on the same class of dedicated hardware you plan to deploy on.