What is rotary position embedding (RoPE)?
Rotary position embedding, RoPE, is a way of telling a transformer where each token sits in a sequence by rotating its query and key vectors before the attention dot product, rather than adding a separate position signal to the token's representation. Each vector gets rotated by an angle proportional to its position in the sequence. Because a dot product between two rotated vectors depends only on the angle between them, the resulting attention score ends up depending on the relative distance between two tokens' positions rather than their absolute positions.
The problem position embeddings solve
Self-attention compares tokens to each other with no built-in sense of order: shuffle the tokens in a sequence and, without some form of position information, the attention scores between any given pair would be unchanged. Early transformers solved this by adding a fixed or learned position vector to each token's embedding before the first layer, injecting absolute position as extra information mixed into the token representation itself.
What RoPE does differently
Rather than adding a position signal to the token embedding, RoPE rotates the query and key vectors directly, splitting each vector into pairs of dimensions and rotating each pair by an angle that grows with the token's position, using a different rotation frequency per pair. The key property this produces is that the dot product between a query at position i and a key at position j comes out as a function of the difference i minus j, not of i and j individually. Two tokens five positions apart produce the same relative signal whether that pair sits at the start of the sequence or near the end, which is a more directly useful signal for attention than raw absolute position.
Why relative position, not absolute, is what attention needs
What a token needs to know about another token, for most language patterns, is how far away it is and in which direction, not which absolute index either one occupies. A verb generally needs to find its subject a few tokens back regardless of whether that pair happens to sit at token 10 or token 10,000. Encoding relative distance directly into the attention score, instead of asking the model to infer it from two absolute position signals, is why RoPE has become the default choice in most current open-weight LLMs, replacing the earlier added-position-vector approach.
How it connects to context length
RoPE's rotation frequencies are set relative to the range of positions seen during training, and running a model well beyond that trained range can degrade its attention quality unless those frequencies are rescaled for the longer range, an adjustment often called RoPE scaling or context extension. This is a different concern from the KV cache memory a longer sequence requires, which our KV cache post covers, and from the context window post's coverage of what a stated context length means for a given model. See handling long context requests for what this looks like from the serving side, including on a Spark's 128GB unified memory, where the practical limit on context length is usually memory for the KV cache rather than the position encoding scheme itself.