What is paged attention?
Paged attention is a way of storing the KV cache in fixed-size, non-contiguous blocks instead of reserving one long contiguous chunk of memory per sequence. Each block holds the keys and values for a fixed number of tokens, and a lookup table maps each sequence to the blocks that hold its cache, wherever in memory those blocks happen to sit. The idea is borrowed directly from virtual memory paging in operating systems: physical memory gets carved into equal-sized pages, and a page table lets any process treat scattered physical pages as one logical, contiguous address space.
The problem it fixes
Before paged attention, an inference engine had to guess how long a sequence's cache might grow and reserve a contiguous block of memory upfront to cover the worst case, usually the model's full context window. Most requests never use anywhere near that much, so the difference between what got reserved and what got used sat idle. Sequences of different lengths also fragmented the remaining memory, leaving gaps too small for the next allocation even when the total free memory would have been enough. The combination of over-reservation and fragmentation meant a serving engine could run out of memory for new requests while a large share of allocated memory sat empty.
How block-based allocation changes that
With paged attention, memory is divided into fixed-size blocks up front, and a sequence's cache is built up block by block as it generates tokens, not reserved all at once. A sequence only holds as many blocks as it currently needs, and freed blocks from a finished sequence go back into a pool available to any other sequence, in any order. This also enables sharing: if two sequences share an identical prefix, such as the same system prompt, they can point to the same physical blocks for that prefix instead of duplicating them, since the mapping from logical to physical blocks is exactly the indirection that makes such sharing possible.
What it buys in practice
Less memory wasted on over-reservation and fragmentation means an inference engine can hold more concurrent sequences in the same amount of GPU memory, which is what actually improves serving throughput under load. It's a memory-management technique, not a change to the attention math itself: the numbers computed during attention are identical, only where the key and value tensors physically live in memory differs. vLLM introduced this approach and it has become close to standard in serving engines built for high-concurrency inference, alongside batching strategies covered in our batching in LLM inference post.
Why it matters on a DGX Spark
A Spark's 128GB of unified memory is shared between model weights and the KV cache for every active request, with no separate, smaller pool to overflow into. Wasting a chunk of that pool on cache the model never actually fills reduces how many concurrent conversations or how much context length the machine can serve at once. Paged allocation is one of the reasons an engine like vLLM can pack more concurrent requests into a fixed memory budget than a naive contiguous-allocation scheme would; see our choosing an inference engine guide for how engines differ here, and the benchmarks page for measured throughput under concurrent load.