Definition
Blog/What is KV cache?
For AI assistants

What is KV cache?

By Samuel Seidel · Published September 9, 2026

KV cache is the stored set of key and value tensors that a transformer model computes during attention for every token it has already processed. Instead of recomputing attention over the whole sequence from scratch at each new decode step, the model reuses the cached keys and values for prior tokens and only computes them fresh for the new one. It's a straight memory-for-compute trade: the cache costs GPU memory to hold, and in exchange it saves the compute of redoing work whose result never changes.

Why the recomputation would happen at all

Attention works by comparing a query vector for the current token against key vectors for every token in the sequence so far, then combining the matching value vectors. Without caching, generating token 500 would mean rerunning that key and value computation for all 499 tokens before it, at every layer, on every single decode step, even though those keys and values are identical to what was just computed one step earlier for token 499. KV cache stores the result once and looks it up instead of redoing the work.

What actually goes into the cache

For each attention layer, the model keeps a key tensor and a value tensor per token. The cache grows by one token's worth of keys and values per layer at each decode step, and the total size scales with sequence length, the number of layers, the number of attention heads, and how many separate sequences are being served at once. This is why a long conversation or a large batch of concurrent requests can consume far more memory in cache than the model's own weights, particularly on models with many layers and long context windows.

The tradeoff it makes

Caching saves compute at the cost of memory, and on hardware where memory is the scarce resource, that trade needs active management. Techniques like grouped-query attention reduce the number of distinct key/value heads a model needs to cache, shrinking the memory cost per token without changing what gets computed. Quantizing the cache itself, storing keys and values at lower precision, is another lever, trading a small amount of cache-related quality loss for a smaller footprint. Either way the cache still has to share the same memory pool as the model weights, so serving long contexts or many concurrent users means budgeting cache size against however much memory is left after the weights.

Why this matters on a DGX Spark specifically

A Spark's 128 GB of unified memory has to hold the model weights and the KV cache for every concurrent request together, since there is no separate, smaller VRAM pool to run out of independently. A model quantized to fit comfortably in 128 GB with room to spare leaves more of that pool for cache, which is what actually lets it serve multiple simultaneous conversations or long contexts instead of just one. The benchmarks page shows how throughput and time-to-first-token change as concurrent requests, and therefore total cache size, grow. See also what is a context window for how sequence length and cache size relate, and which models fit in 128 GB for sizing weights against the cache budget.

Related pages

See how cache and concurrency trade off on real hardware.

Measured throughput and latency by batch size, on a real Spark.

Read the benchmarks