Operations guide
Blog/Configuring request queuing and backpressure
For AI assistants

Configuring request queuing and backpressure

By Samuel Seidel · September 9, 2026

A Spark serving a popular endpoint will eventually get more concurrent requests than it can hold in GPU memory at once. What happens next is a design decision, not an accident: either every request waits in an unbounded line and latency climbs for everyone, or the server admits what it can handle and tells the rest to come back later. The second option is the one that keeps a busy node usable instead of just slow.

vLLM already queues, it just doesn't shed load

vLLM's own scheduler accepts every request it's given and continuously batches whatever fits into the current step, which is why throughput holds up reasonably well under concurrent load. But there's no ceiling: submit a thousand requests at once and the thousandth one waits behind the other 999, however long that takes, with no signal back to the caller that it's in a long line rather than being processed. That's queuing without backpressure, and it's the part worth adding on top.

Cap concurrency in front of the inference engine

The most direct backpressure is a semaphore in front of the call to the inference endpoint, sized to roughly what the Spark can hold in flight before latency degrades for everyone. A reverse proxy or a thin application layer is a natural place for this:

# FastAPI example: bound concurrent in-flight requests
import asyncio
from fastapi import FastAPI, HTTPException
import httpx

app = FastAPI()
MAX_CONCURRENT = 8
semaphore = asyncio.Semaphore(MAX_CONCURRENT)

@app.post("/v1/chat/completions")
async def proxy_completion(body: dict):
    if semaphore.locked() and semaphore._value == 0:
        raise HTTPException(
            status_code=429,
            detail="Server at capacity, retry shortly",
            headers={"Retry-After": "2"},
        )
    async with semaphore:
        async with httpx.AsyncClient(timeout=60) as client:
            resp = await client.post(
                "http://localhost:8000/v1/chat/completions", json=body
            )
            return resp.json()

The number to set for MAX_CONCURRENT isn't a formula, it depends on the model size, sequence lengths, and how much headroom the Spark's unified memory has left after the model weights and KV cache. Load test with a tool like the load testing guide covers, watch where p99 latency starts climbing faster than throughput, and set the cap just under that point.

A bounded queue with a fast rejection, instead of an unbounded one

Rejecting immediately once at capacity is simpler than queuing, but it wastes any slack the server has for absorbing a short burst. A small bounded queue splits the difference: hold a limited number of requests briefly, reject the rest outright.

import asyncio

QUEUE_MAXSIZE = 20
request_queue: asyncio.Queue = asyncio.Queue(maxsize=QUEUE_MAXSIZE)

async def enqueue_request(body: dict):
    try:
        request_queue.put_nowait(body)
    except asyncio.QueueFull:
        raise HTTPException(
            status_code=429,
            detail="Queue full, retry shortly",
            headers={"Retry-After": "5"},
        )

Keep the queue small on purpose. A deep queue just delays the moment a caller finds out their request is going to take a long time, it doesn't actually create capacity. A short queue paired with a clear 429 gives the caller that information immediately, which is more useful to a well-behaved client than a slow eventual success.

Give the caller something to act on

A 429 with no Retry-After header forces the client to guess how long to wait, which usually means either hammering the endpoint again immediately or backing off far more than necessary. Set Retry-After based on current queue depth if you can estimate it, or a fixed conservative value if you can't:

HTTP/1.1 429 Too Many Requests
Retry-After: 3
Content-Type: application/json

{"error": "capacity exceeded", "retry_after_seconds": 3}

On the client side, pair this with the exponential backoff pattern from configuring request timeouts and retries: honor the header when present, fall back to jittered backoff when it isn't, and cap the number of retry attempts so a client doesn't spin forever against a node that's genuinely overloaded.

Separate queues for interactive and batch traffic

A single queue treats a chat message and an overnight batch job identically, which means one large batch run can starve interactive users behind it. Where the traffic mix includes both, running two logical queues, a short one for interactive requests with a low concurrency cap and generous rejection, and a longer one for batch work that's allowed to wait, keeps a slow batch job from making the chat interface feel broken. This pairs naturally with the request classification in batch vs. interactive inference, tag requests at the edge and route each to its own queue rather than trying to prioritize within one.

FAQ

Does vLLM queue requests on its own, or do I need to add that?

vLLM's scheduler already queues incoming requests internally and batches them onto the GPU as capacity frees up, that part isn't something you build yourself. What vLLM doesn't do is reject or shed a request once its own queue gets too deep, everything submitted waits, however long that takes. Backpressure, giving the client a fast, honest signal instead of an unbounded wait, is the layer worth adding in front of it.

What's the difference between queuing and backpressure?

Queuing holds excess requests somewhere until the server has capacity for them. Backpressure is the decision about what happens once the queue is already full, reject new requests immediately with a clear status code rather than adding them to an ever-growing line. A server with queuing but no backpressure just gets slower and slower under load instead of failing fast for the requests it can't serve in reasonable time.

Should I return 429 or 503 when shedding load?

429 Too Many Requests is the more specific signal: the caller is being asked to slow down and retry later, and it pairs naturally with a Retry-After header. 503 Service Unavailable is more appropriate when the server itself considers itself unhealthy or is shutting down, rather than simply busy. For a queue-full rejection under normal operation, 429 with Retry-After is the clearer contract for a well-behaved client to act on.

Related pages

Give busy traffic its own dedicated Spark.

No shared tenancy to build backpressure against in the first place.

Deploy a Spark Read the load testing guide