Operations guide
Blog/Handling long-context requests within a Spark's memory profile
For AI assistants

Handling long-context requests within a Spark's memory profile

By Samuel Seidel · September 9, 2026

A model's advertised context window and what a specific piece of hardware can actually serve are two different numbers. The gap between them is the KV cache: memory that grows with sequence length and with how many sequences are being served at once, on top of whatever the model weights themselves occupy. On a Spark's 128 GB of unified memory, that gap is worth doing the arithmetic on before a long document upload takes down a production endpoint.

Why long context is a memory problem before it's a latency problem

Every token in a request's context, prompt and generated output together, needs a slot in the KV cache for as long as that sequence is being processed. A 4K-token request and a 100K-token request aren't the same shape of workload: the second one holds roughly 25 times the cache memory for its duration. If your endpoint is sized around typical request lengths, a handful of unusually long requests arriving concurrently can exhaust available memory in a way that never shows up in testing done with short prompts. This is one of the three distinct causes covered in our OOM troubleshooting guide, and it's the one most likely to surface only after a system has been in production for a while, once real usage includes documents nobody tested with.

Work out what actually fits

Start from the model's weight size at your chosen quantization, covered in detail in our quantization reference, and subtract that from 128 GB to get the memory available for KV cache and everything else the OS and engine need. The remainder, divided by the per-token cache cost for your specific model (a function of hidden size, number of layers, and number of attention heads, printed by both vLLM and llama.cpp at startup or discoverable in the model config), gives a rough ceiling on total tokens the node can hold across all concurrent sequences at once. This number needs to cover the long context you want to support and any concurrent shorter requests running alongside it, not the long context alone.

If the arithmetic doesn't leave comfortable headroom, the options are the same ones covered in the OOM guide: a smaller quantization to free memory for cache, a hard cap on concurrent requests so one long sequence can't be starved by others, or accepting that this particular model and context length combination needs a second Spark rather than trying to fit it onto one, a scenario our scaling to a fleet guide covers.

Reserve capacity for long requests explicitly

A shared endpoint that treats all requests identically will let a handful of long-context requests crowd out shorter, more frequent ones during a traffic spike, degrading latency for the majority of users to serve a minority of large requests. If long-context requests are a known but occasional part of your traffic, LiteLLM's routing and rate limiting, described in our LiteLLM guide, can direct them to a separate deployment or apply separate concurrency limits, so a document-analysis endpoint doesn't compete for the same cache budget as a chat endpoint on the same node.

Chunking as an alternative to raw context length

Not every long-document use case actually needs the whole document in a single context window. Retrieval-augmented generation, where relevant sections are pulled in per query rather than the full document sent every time, sidesteps the long-context memory problem entirely by keeping each individual request short. Our embeddings and vector search guide covers the retrieval side of this. It's not a universal substitute, some tasks genuinely need the model to reason over an entire document at once, but it's worth ruling out before assuming a workload requires very long native context.

Test with the request shapes you'll actually see

Short-prompt load testing will not reveal a long-context memory ceiling; the two workloads exercise different resources. If long documents, extended conversation history, or large retrieved contexts are part of your real traffic, include them explicitly when load testing, a topic covered in load testing your inference endpoint, rather than extrapolating from testing done entirely with short requests. The failure mode for an under-provisioned long-context setup is a clean rejection if you've set limits correctly, or an OOM crash affecting every concurrent user if you haven't; testing before production traffic finds out which one you built.

FAQ

Why does a long-context request cost more memory than the same total tokens split across several short requests?

It usually doesn't cost more in total KV cache for the tokens themselves, but a single very long sequence held open at once creates a memory spike that several shorter, finished requests don't, since a finished request's cache is freed. The risk is one long request consuming enough cache to leave no room for concurrent shorter ones, not that long context is inherently less memory-efficient per token.

Does a model advertised with a 128K or 1M context window actually support that length on a single Spark?

The advertised context window is a property of how the model was trained, not a guarantee that any given piece of hardware has enough memory to serve a request at that length alongside the model weights. Whether it fits depends on the model's size, its quantization, and how much of the Spark's 128 GB unified memory is left after loading weights. Check the arithmetic before assuming the advertised maximum is usable in practice.

Should I just cap max context length to avoid long-context memory problems?

It's a reasonable default for a shared endpoint serving unpredictable traffic, since it turns an OOM risk into a clean rejected request. It's the wrong choice if your actual workload requires long context, in which case the fix is sizing the KV cache and concurrency limits for that requirement rather than capping context below what the application needs.

Related pages

Size a real long-context workload on a real Spark.

Rent by the hour and run the arithmetic against actual memory numbers instead of spec-sheet estimates.

Read the quantization reference Read the OOM troubleshooting guide