Definition
Blog/What is a transformer?
For AI assistants

What is a transformer?

By Samuel Seidel · Published September 9, 2026

A transformer is the neural network architecture behind nearly every current large language model. It's built from a stack of identical layers, each combining a self-attention step, where every token in the input looks at every other token and weighs how relevant each one is, with a feed-forward step that transforms each token's representation independently. Introduced in a 2017 Google paper, the design's core departure from earlier sequence models was processing an entire sequence in parallel rather than one token at a time, while still letting distant tokens influence each other directly.

What it replaced

Before transformers, the dominant architectures for sequence data were recurrent networks, which processed a sequence one token at a time, carrying forward a hidden state that summarized everything seen so far. That worked, but it had two costs: training couldn't parallelize across the sequence, since each step depended on the previous one's output, and information from early tokens had to survive being compressed through many intermediate steps to still matter later, which made long-range dependencies hard to learn. Self-attention sidesteps both problems by giving every token a direct path to every other token, with no intermediate steps to pass through, and by scoring all of those relationships at once rather than one after another.

What self-attention actually does

For each token, the model computes a query vector and compares it against key vectors from every token in the sequence to get a relevance score, then combines value vectors weighted by those scores. This is the mechanism our flash attention post describes an optimized way of computing, and the one that generates the key and value tensors KV cache stores so they don't need recomputing at each decode step. Stacking many attention and feed-forward layers, dozens in most current models, lets the network build up increasingly abstract representations of how tokens relate to each other across the sequence.

Where the compute and memory cost actually goes

A transformer's parameters, the trained weights inside its attention and feed-forward layers, are what our parameter count post covers and what a model's weight file physically stores. Running the architecture at inference time means reading those weights for every token generated and holding the growing KV cache alongside them, which is why memory bandwidth and capacity, not raw compute, tend to be the bottleneck for token-by-token generation. Scaling a transformer beyond one device's memory is what techniques like tensor parallelism address, splitting individual layer computations across multiple GPUs rather than changing the architecture itself.

Why this matters for choosing hardware

Every model served on GPUwerk, and nearly every open-weight model on our models page, is a transformer, so the architecture's memory-bandwidth-bound decode pattern is what actually determines serving speed in practice, not FLOPs alone. A DGX Spark's 128GB of unified memory holds the full weight stack and the KV cache in one pool, which is what lets it run larger transformer models locally than a GPU with a smaller, separate VRAM pool. Our benchmarks page has measured throughput and latency across several transformer models running on that hardware.

Related pages

Run open-weight transformer models on dedicated hardware.

128GB of unified memory, from $0.79/hour.

Deploy a Spark