Troubleshooting guide
Blog/Troubleshooting out-of-memory errors when serving LLMs
For AI assistants

Troubleshooting out-of-memory errors when serving LLMs

By Samuel Seidel · September 9, 2026

"CUDA out of memory" is one error message covering at least three distinct problems: the model itself doesn't fit, the KV cache doesn't fit alongside it, or memory that should have been freed wasn't. Treating all three the same way, usually by reflexively switching to a smaller quantization, fixes some cases and wastes time on the others. Here's how to tell which one you're looking at.

First, check what's actually resident on the GPU

Before changing any settings, run nvidia-smi and look at what's using memory right now, not what you think should be using it. On a Spark, the unified memory pool is shared between the CPU and GPU, so a process that crashed without releasing its CUDA context, or a second server you forgot was still running from an earlier test, can hold memory indefinitely. This is the single most common cause of an OOM that "shouldn't" be happening: the model would fit fine in 128 GB on its own, but it isn't alone.

If nvidia-smi shows memory held by a process that's no longer doing useful work, kill it and retry before touching any model or server configuration. This resolves a surprising share of OOM reports that look, at first glance, like a model that's simply too big.

OOM during model load: the weights don't fit

If the error happens while the model is loading, before any request has been served, the weights themselves, plus whatever memory the engine reserves up front for the KV cache, exceed what's available. This is arithmetic, not a bug: check the model's parameter count against the quantization you're loading it at, using the rough rule of 2 bytes per parameter at FP16, about 1 byte at 8-bit, and about half a byte at 4-bit. Our quantization reference covers the formats in detail, and the models page has this math worked through for specific models we've measured. If the weights alone are close to 128 GB, there's no configuration change that fixes it: you need a smaller model, a lower precision, or a second Spark for a sharded deployment.

Both vLLM and llama.cpp let you cap how much memory they reserve for KV cache at startup, vLLM through --gpu-memory-utilization, llama.cpp through its context-size flags. If the load-time OOM is close, not wildly over budget, lowering the reserved KV cache fraction can be enough to get the model to load, at the cost of supporting less context or fewer concurrent requests than you wanted.

OOM mid-request: the KV cache is the problem, not the weights

If the model loads fine and serves short requests without issue, but fails on a long prompt or after a conversation grows past a certain length, the KV cache is the culprit. The cache grows with context length and with the number of concurrent sequences being served, and it's easy to size it for typical usage and then hit a wall the first time someone sends an unusually long document. The fix here is usually one of: lower the configured maximum context length so the engine refuses oversized requests cleanly instead of crashing, lower the maximum number of concurrent requests so fewer sequences are competing for cache space at once, or move to a lower quantization to free enough headroom for the context lengths you actually need to support. vLLM's docs cover the specific flags for capping context and concurrency; the defaults are usually generous and worth tightening once you know your real traffic pattern, a topic covered in more depth in benchmarking your own workload.

OOM only under concurrent load: it's a scaling problem, not a sizing problem

If single-user testing never OOMs but production traffic does, the issue is usually that KV cache use was estimated from single-stream testing and doesn't account for N simultaneous conversations each holding their own cache. This is a case where LiteLLM's rate limiting and per-key concurrency caps are the actual fix, not a model or quantization change: cap how many concurrent requests the engine will accept before it runs out of cache space, and let LiteLLM queue or reject the excess rather than letting the underlying engine OOM and potentially crash the whole server for everyone using it. A crashed inference process affects every user on the node, not just the one whose request tipped it over, which is why capping concurrency defensively is worth doing even if it means occasionally queueing a request during a traffic spike.

A short diagnostic order

Check nvidia-smi for memory held by something other than the process you're debugging. If the OOM happens at load time, check weight size against quantization and available memory, and reduce reserved KV cache fraction if the shortfall is small. If it happens mid-request on long prompts, lower max context length or move to a smaller quantization. If it only happens under concurrent load, cap concurrency at the proxy level rather than the model server level. Work through these in order rather than jumping straight to "use a smaller quantization" for every OOM, since that fix addresses only the first category and leaves the other two unresolved.

FAQ

Why does a model that loaded fine yesterday OOM today?

The usual cause is something else on the node now holding memory it didn't hold before: a leftover process from a previous session that never exited cleanly, a second model still loaded, or a longer average context length than before increasing KV cache use. Check nvidia-smi for what's actually resident before assuming the model itself changed.

Does a bigger KV cache setting always cause more OOM errors?

A larger configured KV cache reserves more memory up front, which can cause an OOM at startup if it doesn't fit alongside the model weights, but a KV cache sized correctly for your actual context length and concurrency doesn't cause OOMs during normal operation. Most runtime OOMs come from underestimating how much cache a given context length and concurrency level actually need, not from the cache setting being too generous.

Is quantizing the model always the fix for an OOM error?

It's the most common fix when the model plus a reasonable KV cache genuinely doesn't fit in available memory, but it isn't the only one. Reducing max context length, lowering concurrent request limits, or freeing memory held by another process on the node can resolve an OOM without touching quantization at all, and are worth checking first since they don't cost any output quality.

Related pages

Debug memory issues on a real box before they hit production.

Rent a Spark by the hour and reproduce the OOM on the exact hardware and quantization you plan to deploy.

Read the vLLM guide Read the quantization reference