What is flash attention?
Flash attention is a way of computing the attention operation inside a transformer that avoids ever writing the full attention matrix out to GPU memory. Standard attention computes a score between every pair of tokens in a sequence, producing a matrix that grows with the square of sequence length, then writes that matrix to memory, reads it back for the softmax step, and reads it again to combine it with the value vectors. Flash attention fuses these steps into one kernel and processes the sequence in tiles, keeping intermediate results in the GPU's fast on-chip memory and only ever writing the final output. The math it computes is identical to standard attention; what changes is how much data moves in and out of memory to get there.
Why memory traffic, not compute, is the bottleneck
The naive read on attention is that it's expensive because of the number of multiplications involved, and at long context lengths that's part of it. But on modern GPUs, moving data between the GPU's main memory and its on-chip compute units is often slower than the arithmetic itself. Standard attention's problem is that it writes the full N-by-N score matrix to memory and reads it back multiple times, and that matrix gets large fast: double the sequence length and it quadruples. Flash attention restructures the computation so those intermediate values never leave fast on-chip memory in the first place, which cuts memory traffic even though the number of floating-point operations stays roughly the same.
What it changes in practice
Two things follow directly from lower memory traffic. First, memory use during attention drops from scaling quadratically with sequence length to scaling linearly, which is what lets a model handle a longer context window in the same amount of memory that used to cap out much earlier. Second, since attention spends less time waiting on memory, the attention portion of a forward pass runs faster, particularly during prefill, when the whole prompt is processed at once and the attention matrix would otherwise be largest. Our benchmarks page shows prefill on the Spark running around 1,900 to 1,950 tokens per second, in the compute-bound regime where a kernel like this earns its keep; decode is bounded by memory bandwidth for reading model weights, a separate bottleneck that flash attention doesn't touch.
Why longer context needs it
Without a fused kernel, doubling context length doesn't just double the work, it roughly quadruples the memory the attention matrix itself needs. That's the practical wall that made very long context windows impractical before kernels like this became standard in inference engines. See our handling long context requests post for what that looks like from the serving side, and our what is a context window post for the concept itself.
Where it fits in a typical inference stack
Flash attention is implemented as a kernel inside inference engines like vLLM and llama.cpp, not something you configure separately in most setups; it's usually on by default when the hardware and model support it. It's one of several optimizations, alongside things like KV caching, that determine how much context and how many concurrent requests a given amount of memory can actually serve. See our choosing an inference engine guide for how engines differ on this.