Foundational Mechanics · 7 min
Why order needs to be encoded
Self-attention is blind to word order by construction; four positional schemes put order back in, and each one breaks differently past its training length.
Take "the dog bit the man" and scramble it to "man the the dog bit." To you these are two different sentences, one alarming and one gibberish. To the core of a transformer — self-attention on its own — they are the same input. That is not a rough edge you sand off later. It is a structural fact about how attention works, and every model you have ever used has to actively fight it. This lesson is about the fight.
The mental model: attention is a set operation
Self-attention computes, for each token, a weighted average over every other token. Strip the notation down and it is:
Attention(Q, K, V) = softmax(Q Kᵀ / √d) · V
Q, K, and V are the query, key, and value matrices; each row is one token's vector. Look for a row's index in that formula and you will not find it. Token 3 is treated exactly like token 30. Shuffle the input rows and every dot product qᵢ · kⱼ is still computed, just relabelled, and the outputs come out shuffled the same way. The technical name for this is permutation-equivariant: permute the input, and the output permutes identically, with nothing else changed.
So a bare attention stack literally cannot tell "dog bit man" from "man bit dog." The set of token vectors is identical; only the order differs, and order is the one thing the operation discards. One caveat, so a sharp reader can't catch you out: this is exactly true of the attention operation on its own. A decoder-only model also imposes a causal mask — each token may attend only to the ones before it — and that mask is not symmetric under permutation, so it quietly breaks the tie; research has shown that causal Transformers can extract some positional information even with no explicit positional encoding at all. The schemes below make that signal explicit and controllable instead of leaving it an accident of the mask. In English word order carries most of the grammar, which makes a position-blind model close to useless. Position has to be put back in by hand. The four schemes below are four ways to do it.
Mechanics: four ways to inject position
Absolute sinusoidal (Vaswani et al., 2017). The original transformer added a fixed vector to each token embedding, one vector per position, built from sines and cosines at geometrically spaced frequencies. Position 0 gets one pattern, position 1 another. Because the frequencies span fast to very slow, every position gets a unique, smoothly varying signature. No parameters, and in principle it is defined for any position you ask about.
Learned absolute (GPT-2, BERT). Swap the fixed sinusoids for a trainable lookup table: row i is "the position-i vector," learned by gradient descent. Simple and effective, with one hard wall. Train with 1024 rows and there is no row 1025. Ask about token 2048 and there is nothing to add — the table has no entry for it. That is precisely why GPT-2's context was pinned at 1024 and could not be stretched.
Rotary (RoPE — Su et al., 2021). Instead of adding a position vector, RoPE rotates the query and key vectors by an angle set by their position. Group the dimensions into pairs, treat each pair as a point in a plane, and spin it by m·θ for a token at position m, where θ differs per pair (fast spin for some, glacial for others, on the same base-10000 frequency ladder the sinusoids use).
The payoff falls straight out of trigonometry. Take the dot product of a query at position m with a key at position n, both rotated, and the result depends only on the difference m − n, never on m and n separately:
in: rotate q by m·θ, rotate k by n·θ
out: score depends only on (m − n)
q(pos 5) · k(pos 3) ≡ q(pos 105) · k(pos 103)
same relative offset → same scoreRoPE encodes absolute position going in, but the model experiences it as relative distance, which is usually what matters ("the adjective two words back," wherever that pair sits in the sentence). That is why it became the default: Llama, Qwen, and Mistral all run on it.
ALiBi (Press et al., 2021). The most stripped-down of the four. Nothing added, nothing rotated. ALiBi subtracts a penalty from each attention score before the softmax, sized by how far apart the two tokens are:
score(i, j) = qᵢ · kⱼ − m · (i − j)
(i − j) is the distance and m is a fixed per-head slope. Different heads get different slopes — a geometric sequence, for an 8-head model 1/2, 1/4, 1/8, on down to 1/256 — so some heads stay sharply local while others tolerate long range. The result is a recency bias baked directly into the scores: nearby tokens are cheap to attend to, distant ones pay a linearly growing toll.
Failure modes: what happens past the training length
Feed a model more tokens than it trained on and the positional scheme, more than anything else, decides what breaks.
Learned absolute embeddings do not extrapolate at all. There is no parameter for an unseen position, so you hit the wall described above — a cliff, not a slope.
RoPE degrades instead of cliff-edging, but degrade it does. Past the trained length the low-frequency dimension pairs reach rotation angles the model never saw in training, attention gets noisy, and you watch perplexity climb along with "lost in the middle" behavior. The standard fix is to rescale the rotation frequencies so a long sequence looks like a familiar range. Position Interpolation (Chen et al., 2023) squeezes positions linearly; YaRN (Peng et al., 2023) does it per frequency band. RoPE scaling and related long-context techniques have been behind many of these extensions.
ALiBi was built for extrapolation, and that is its headline. Press et al. trained a 1.3B model on length-1024 sequences and ran inference at 2048, matching the perplexity of a sinusoidal model trained at 2048, while training 11% faster and using 11% less memory. The linear penalty is defined at any distance, so nothing snaps when the sequence grows. The catch: the same recency bias that lets it extrapolate also makes it structurally reluctant to reach far, which hurts tasks that need a genuine long-range lookup.
Builder: Before you fine-tune for longer context, learn which scheme the base model uses. RoPE means you are reaching for a scaling factor (PI or YaRN) plus some long-sequence fine-tuning. Learned-absolute means you cannot widen the window at all without surgery on the embedding table. Read the config before you promise anyone 128k.
Defender: Positional handling is an attack surface. An injection payload buried in the middle of a long document exploits exactly the mid-context weakness that sloppy RoPE scaling creates — the model attends poorly there, so a planted instruction behaves unpredictably. When you red-team a model at its advertised max context, test the middle of the window, not just the ends.
Researcher: RoPE's relative-distance property and ALiBi's recency bias are different inductive biases, not just different code. If you are probing length generalization, hold data and architecture fixed and swap only the positional scheme. It is one of the cleanest ablations in the transformer, and the effect size is large.
The through-line: attention hands you a powerful content-mixing engine that is, by itself, blind to order. Grammar, long documents, the exact place a model falls apart past its training length — all of it traces back to the small choice of how you told the model what came first. Keep this failure mode distinct from the two others you will meet later: running out of memory for the KV cache (the resource ceiling of the KV-cache module) and attention diluting itself across a crowded window (the "context rot" of the context-length module) are separate problems with separate fixes. The one this lesson owns is positional — the model meeting rotation angles or position indices it never trained on. When we reach context-length extension and "lost in the middle" in that later module, RoPE's frequency ladder is the set of dials you will be turning.
Sources
- Vaswani et al. (2017), Attention Is All You Need — sinusoidal absolute positional encoding (arXiv:1706.03762).
- Su, Lu, Pan, Wen, Liu (2021), RoFormer: Enhanced Transformer with Rotary Position Embedding (arXiv:2104.09864).
- Press, Smith, Lewis (2021), Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation (arXiv:2108.12409).
- Chen et al. (2023), Extending Context Window of Large Language Models via Positional Interpolation (arXiv:2306.15595).
- Peng et al. (2023), YaRN: Efficient Context Window Extension of Large Language Models (arXiv:2309.00071).
- EleutherAI, Rotary Embeddings: A Relative Revolution — intuitive RoPE walkthrough (blog.eleuther.ai/rotary-embeddings). — https://blog.eleuther.ai/rotary-embeddings/