Streaming vs non-streaming responses for self-hosted LLMs
Both vLLM and llama.cpp serve an OpenAI-compatible API with a stream parameter that switches a completion request between two delivery modes. The model generates tokens at the same rate either way; what changes is whether the client gets them one at a time as they're produced or waits for the whole response to finish. That distinction matters for perceived latency, error handling, and how the reverse proxy and any middleboxes in front of the endpoint need to be configured.
What actually changes between the two modes
In non-streaming mode, the client sends a request and receives nothing back until generation is complete, then gets the full text in one response. In streaming mode, the client receives each token, or small batches of tokens, as server-sent events as soon as the engine produces them, and assembles the full text on its own side as chunks arrive. Total time to the last token is roughly the same in both modes; the difference is time to the first visible token, which is near-instant with streaming and equal to the full generation time with non-streaming. For a chat interface where a person is watching the response appear, that difference is the entire point.
When streaming is worth the added complexity
Any interactive use case where a person is waiting on the response benefits from streaming: chat interfaces, coding assistants, anything where perceived responsiveness matters more than the exact wall-clock time to completion. A 30-second generation feels dramatically different when text starts appearing after half a second versus when the screen stays blank for the full 30 seconds, even though the underlying computation took the same time either way.
For a backend job with no one watching in real time, streaming adds client-side complexity, handling partial chunks, reassembling text, dealing with a dropped connection mid-stream, without a corresponding benefit. Batch classification, embedding generation, or a scheduled report that gets written to a database when done are all better served by non-streaming, which is simpler to implement correctly and just as fast in aggregate. Our batch processing guide covers this pattern in more depth.
Error handling is genuinely different between the two modes
Non-streaming gives you one clean failure mode: the request either returns a complete response or it returns an error, and there's no ambiguous middle state to handle. Streaming can fail partway through, after some tokens have already reached the client, which means the client needs logic for a stream that stops mid-response: was that the model finishing normally, a network drop, or a server-side error partway through generation? The OpenAI-compatible API signals normal completion with a distinct terminal event, but an abrupt connection drop looks different from that and needs to be handled as its own case, typically by retrying the whole request rather than trying to resume a partial stream.
Proxy and load balancer configuration
Streaming responses need to actually stream through whatever sits between the client and the inference engine. A reverse proxy configured to buffer the full response before forwarding it, which is nginx's default behavior, turns a streaming response back into an effectively non-streaming one from the client's perspective, quietly defeating the reason you turned streaming on in the first place. Our guide on setting up TLS for your inference endpoint covers disabling proxy buffering specifically for this reason. If you're running LiteLLM in front of multiple engines, confirm it's passing streamed responses through rather than aggregating them, since a proxy layer added later in a project's life is a common place for this to break silently.
Concurrency and connection lifetime
A streaming connection stays open for the full duration of generation, which is longer than a non-streaming connection stays open for the same request in wall-clock terms, since non-streaming's connection is only held during the network round-trip of the final response, not the generation itself, from the proxy's point of view, depending on how the proxy is configured. In practice this means a server handling many concurrent streaming connections needs a request handling model built for long-lived connections rather than short request-response cycles; both vLLM and LiteLLM are async and built for this, but a naive client-side implementation using a synchronous HTTP client per request can exhaust connection pools faster than expected under load. Worth testing explicitly as part of the load testing described in load testing your inference endpoint, since streaming and non-streaming can behave differently under the same concurrent request count.
Picking a default
For a general-purpose endpoint serving both interactive and background use cases, exposing both modes through the standard stream parameter, as both engines already do, and letting each client choose based on its own use case is simpler than trying to pick one default for everyone. If you have to pick one default for a single-purpose endpoint, match it to who's consuming it: streaming for anything with a human watching, non-streaming for anything automated.
FAQ
Does streaming make the model generate tokens faster?
No. The generation rate, tokens per second, is the same either way; streaming only changes when the client sees each token. Non-streaming waits for the entire completion before sending anything back, so the client perceives the full generation time as a single wait, while streaming sends each token as it's produced.
Is streaming harder to implement than non-streaming?
On the server side, both vLLM and llama.cpp support streaming out of the box through the same OpenAI-compatible endpoint, so there's little extra setup. The added complexity is on the client, which needs to parse a stream of partial chunks and handle a connection that can drop mid-response, instead of making one request and getting one response back.
Should batch or background jobs use streaming?
Usually not. Streaming's benefit is reducing perceived latency for something a person is watching in real time. A batch job with no one watching gets no benefit from incremental delivery and picks up the added complexity of handling a stream for no reason, so non-streaming is the simpler and equally fast choice there.