Operations guide
Blog/Configuring request timeouts and retries for an inference endpoint
For AI assistants

Configuring request timeouts and retries for an inference endpoint

By Samuel Seidel · September 9, 2026

A web API has a fairly predictable response time, and a timeout copied from a REST service's defaults tends to work fine everywhere else. An LLM endpoint doesn't behave that way. A short completion might return in under a second; a long generation against a queued batch can legitimately run past a minute. Get the timeout wrong in either direction and you either cut off real work or leave callers hanging on a request that already failed.

Why LLM latency breaks generic defaults

Request duration on a vLLM or llama.cpp endpoint scales with output length, queue depth, and whatever else is running on the node, see batch vs. interactive inference for how those interact. A 200-token answer and a 4,000-token one aren't the same request with a different number attached; they're different latency classes. A single fixed timeout tuned for the short case kills the long one, and a timeout generous enough for the long case leaves a caller waiting far too long to find out a genuinely stuck request failed.

nginx timeout config for a streaming endpoint

If you're running the reverse proxy setup from setting up a reverse proxy for your Spark, the relevant directive is proxy_read_timeout, which resets on each chunk of data received from the upstream rather than timing the whole request. That makes it a reasonable fit for a streaming completion, as long as you also cap total request duration separately so a stalled stream doesn't hang forever.

# /etc/nginx/conf.d/inference.conf
location /v1/ {
    proxy_pass http://127.0.0.1:8000;
    proxy_http_version 1.1;

    # resets per chunk received, so a slow stream doesn't trip it
    proxy_read_timeout 120s;
    proxy_send_timeout 30s;
    proxy_connect_timeout 5s;

    # disable buffering so tokens reach the client as they're generated
    proxy_buffering off;
    proxy_cache off;
}

120 seconds is a starting point, not a rule. Set it against your actual observed generation times for the longest requests you expect to serve, and if you're running structured output or function calling with retries baked into the generation itself, budget for that too.

Separate the connect timeout from the read timeout

proxy_connect_timeout should stay short, five seconds is generous, because a connection failure to a healthy upstream is either instant or it's a sign the service is down. Conflating it with the read timeout means a genuinely dead backend takes two minutes to report as dead instead of five seconds, which delays both the caller's error and, if you're alerting on it per setting up Prometheus alerts, the alert itself.

Retry only what's safe to retry

A connection refused or a 503 before generation starts is safe to retry: nothing has been generated, so a retry costs a little latency and no correctness risk. A request that fails partway through a stream, or that times out after the model has already produced most of an answer, is not safe to retry blindly, retrying it can double the GPU cost of that call and, if a caller acts on partial output, produce duplicate downstream effects.

# nginx: retry a connection-level failure against the same upstream once, fast
location /v1/ {
    proxy_pass http://127.0.0.1:8000;
    proxy_next_upstream error timeout http_503;
    proxy_next_upstream_tries 2;
    proxy_next_upstream_timeout 5s;
}

Leave application-level retry logic, retrying after a timeout the proxy already let through, to the client. It's the only layer that knows whether re-sending that specific prompt is worth the cost.

Client-side retry with backoff

A minimal retry wrapper around an OpenAI-compatible client, scoped to the failure modes that are actually safe:

import time
import httpx

def call_with_retry(client, payload, max_attempts=3):
    for attempt in range(max_attempts):
        try:
            return client.post("/v1/chat/completions", json=payload, timeout=90)
        except (httpx.ConnectError, httpx.ReadTimeout) as e:
            if attempt == max_attempts - 1:
                raise
            # backoff before retrying against a possibly-overloaded node
            time.sleep(2 ** attempt)

Three attempts with exponential backoff is a reasonable default for an internal service. Don't retry indefinitely; a caller that keeps re-sending a request against a node that's actually overloaded makes the queue worse, not better, and the health checks from writing health checks for your inference endpoint are the better tool for deciding when to route around a bad node entirely rather than retry into it.

FAQ

Why does the default nginx timeout break LLM streaming?

nginx's default proxy_read_timeout is 60 seconds, and it resets on each byte received from the upstream, not on the whole request. A streaming response that keeps sending tokens won't hit it. A request stuck waiting behind a full batch queue with no bytes flowing yet, or a long generation with gaps between tokens under load, can. Raise it explicitly and pair it with a client-side timeout that matches what your callers actually expect to wait.

Is it safe to retry a failed inference request automatically?

Only for requests that failed before generation started, a connection refused, a 503 from an overloaded queue, a timeout on the initial response. Retrying a request that failed partway through a stream risks double-billing the same generation cost and, for non-idempotent side effects downstream of the response, duplicate action. Scope retries to idempotent failure modes and cap the attempt count.

Should retries happen at the proxy or the client?

Both, for different failure classes. The proxy is well placed to retry a connection-level failure against a healthy upstream instantly, before the caller even notices. Application-level failures, a malformed response, a semantic timeout the proxy can't see, need the client to decide whether retrying makes sense for that specific call.

Related pages

Run the proxy on hardware sized for your own traffic.

A dedicated Spark, so a timeout you tune is tuned against your own queue, not a shared tenant's.

Read the reverse proxy guide Read the load testing guide