What is grouped-query attention?
Grouped-query attention, GQA, is an attention design where several query heads share one set of key and value heads, instead of every query head having its own dedicated key and value heads. In standard multi-head attention, a model with 32 attention heads computes and stores 32 separate sets of keys and values. With GQA, those 32 query heads might be split into 8 groups of 4, and each group shares a single key/value head, so only 8 sets of keys and values get computed and cached rather than 32.
Why the number of key/value heads matters
The KV cache stores a key and value tensor per attention head, per token, per layer, and its size scales directly with how many distinct key/value heads a model has. Query heads don't add to that cost the same way, since queries aren't cached across decode steps, only keys and values are. Cutting the number of distinct key/value heads from 32 to 8, while keeping all 32 query heads, shrinks the cache by the same factor without touching the number of query heads that give the model its representational capacity.
Where it sits between two extremes
Standard multi-head attention gives every query head its own key/value head, maximizing representational flexibility at the cost of a large cache. Multi-query attention, an earlier alternative, pushes to the opposite extreme, one shared key/value head for all query heads, minimizing cache size but reducing how differently each head can attend to the sequence. Grouped-query attention is a tunable middle point: more groups keeps more of multi-head attention's flexibility, fewer groups keeps more of multi-query attention's memory savings. Most current open-weight models settle on a small number of groups, trading a modest amount of representational capacity for a substantial cut in cache size.
What it costs, if anything
Because it reduces representational capacity in the key/value projections, GQA in principle could hurt output quality compared to full multi-head attention at the same parameter count. In practice, models trained with GQA from scratch, rather than converted afterward, show little measurable quality difference against multi-head attention at the same overall size, which is why it has become the default in most current LLM architectures rather than a niche efficiency trick. This is distinct from techniques like flash attention, which changes how the attention computation is scheduled on a GPU without changing the model's architecture, and from paged attention, which changes how the resulting cache is laid out in memory; GQA changes how much there is to cache in the first place.
Why this matters for serving on a Spark
A smaller KV cache per token means more concurrent sequences or longer contexts fit in the same amount of GPU memory. On a DGX Spark, where the KV cache shares the full 128GB unified memory pool with the model weights rather than a separate VRAM allocation, a model built with grouped-query attention leaves noticeably more of that pool available for cache than an equivalent model using full multi-head attention. Most current open-weight models on our models page use GQA already, which is part of why they serve the concurrency levels shown on our benchmarks page.