Foundational Mechanics · 6 min

Self-attention, actually explained

Self-attention is a soft, differentiable dictionary lookup where every key matches a little — the source of both its power and its quadratic cost.

Every token in a transformer gets to look at every other token and decide what's worth pulling in. That's the whole trick. The mechanism that implements "look around and decide" is self-attention, and once you see it as a soft, learned lookup, the rest of the architecture stops being mysterious.

The mental model: a lookup where every match is partial

Start with a Python dict. You hand it a query, it compares that against its keys, and on an exact match it returns the one corresponding value:

store = {k1: v1, k2: v2, k3: v3}
result = store[query]        # hard match: exactly one key wins

Self-attention is that same operation made soft and differentiable. No single key wins outright. Every key matches to some degree, and you get back a blend of all the values, weighted by how well each key matched. Nothing is retrieved exactly; everything is retrieved a little.

Three roles do the work, and all three are projected from the token's own vector by three learned weight matrices:

  • Query (Q): what this token is looking for. "I'm a pronoun; which noun do I refer to?"
  • Key (K): what a token advertises about itself. "I'm a singular animate noun, right here."
  • Value (V): the actual content a token hands over once it's been selected.

Q and K are the matchmaking channel; V is the payload. They're separate projections on purpose, so a token can advertise one thing through its key and contribute something else through its value, the way a search engine can match on a title but return the whole document.

The mechanics: one equation

Vaswani et al. (2017) write the entire operation as scaled dot-product attention:

Attention(Q, K, V) = softmax( QKᵀ / √d_k ) V

Read it left to right. QKᵀ dots every query against every key, producing an n×n grid of raw compatibility scores for a sequence of length n. Divide by √d_k, where d_k is the key dimension (64 in the base model). Softmax each row so the scores across all keys turn into weights that sum to 1. Multiply by V, and each token's output becomes a weighted average over every value in the sequence.

The √d_k term earns its place. Dot products of two random d_k-dimensional vectors have variance proportional to d_k, so at d_k=64 the raw scores drift into the tens or hundreds. Push numbers that large into softmax and it saturates: one weight snaps to ~1, the rest collapse to ~0, and the gradient through softmax dies. Dividing by √d_k (which is 8 here) pulls the scores back toward unit variance and keeps gradients alive. Drop it and training on wide heads gets visibly unstable.

A worked micro-example

Three tokens, keys living in a 2-D space (d_k=2, so √d_k ≈ 1.414). We're computing the output for the word "bank", whose query vector is q = [0, 2].

token    key k        q·k    score = q·k / √2
-----    ---------    ---    ----------------
the      [1, 0]        0        0.00
river    [0, 2]        4        2.83
bank     [1, 1]        2        1.41

Softmax over [0.00, 2.83, 1.41]:

exp:     1.00   16.94   4.11      (sum = 22.05)
weight:  0.045  0.769   0.186

"bank" attends 0.77 to "river," 0.19 to itself, 0.05 to "the." Its output is 0.045·v_the + 0.769·v_river + 0.186·v_bank, a vector now dominated by whatever "river" carries. That is disambiguation happening mechanically. The representation of bank has bent toward the river sense because the river token's key matched its query. No rule fired anywhere. The weight matrices that produced these Q and K vectors were trained so this kind of pull shows up exactly where it helps prediction.

Stack this across layers and the mixing compounds. Layer 1's output for "bank" already carries river-flavor, so when layer 2 attends over those blended vectors, information travels several hops in a handful of layers. This is the thing RNNs couldn't do cheaply: the path between any two tokens is one attention hop, O(1), not O(n) steps of a recurrence unrolled across the sequence.

The failure mode: it costs n²

That QKᵀ grid is n×n. Every token scores against every other token, so the attention computation — building that grid and applying it — scales with the square of sequence length. Be precise about what's quadratic here: it's this computation, not the whole model's cost story. The persistent memory a long context occupies is a separate, linear cost, the KV cache, and Lesson 7 (KV caching) picks up that thread to show the linear term often saturates VRAM and memory bandwidth before this quadratic compute becomes the binding constraint. Keep the two costs mentally distinct. The paper's Table 1 puts self-attention at O(n²·d) per layer against a recurrent layer's O(n·d²).

The consequence is arithmetic, and it doesn't negotiate:

context n      score entries (n²)
--------      ------------------
   1,000               1,000,000
   4,000              16,000,000
  32,000           1,024,000,000
 128,000          16,384,000,000

Double the context, quadruple the attention cost. Go from 1k to 128k and the score matrix is roughly 16,000 times bigger. That single fact explains a lot of downstream behavior: why long-context pricing climbs the way it does, why the first token of a long prompt is slow to arrive (prefill has to build the whole triangle before decoding starts), and why "just paste the entire repo" hits a wall that has nothing to do with model quality. The crossover with recurrence sits near n ≈ d. Below your model dimension the n² is cheap; well above it, that term swallows everything.

Builder. Your latency and memory are governed by prefill, and prefill is O(n²). Trimming an 8k prompt to 4k doesn't halve the cost, it quarters the attention portion. Retrieve and prune hard; a tight 2k prompt beats a lazy 20k one on speed and on the bill. The KV cache (the keys and values kept around so each new token skips recomputing the past) grows only linearly per token, but every new step reads it against the entire prior context, which is the other place your memory goes — Lesson 7 is where that story gets told in full.
Researcher. Almost every "efficient attention" line, whether sparse, low-rank, linear, or sliding-window, is an attack on that n² term, trading exact all-pairs mixing for an approximation that scales better. Knowing precisely what the full softmax computes tells you exactly what each approximation gives up.

The security angle

Softmax is a competition with no notion of authority. System prompt, retrieved document, user message, tool output: once tokenized, they're all just rows in the same K and V matrices, competing for weight in the same softmax. Role markers and post-training tilt that competition toward your system prompt, but only as a learned bias, and nothing in the attention math fences off a genuine instruction — so a token reading "ignore previous instructions" can still capture attention mass from one if its key happens to match the model's queries well. There's no architectural boundary here the way an OS separates kernel space from user space. That flat competition is the substrate prompt injection exploits, which is why injection isn't a bug you patch inside the attention layer. The layer is doing exactly what it was built to do. Defenses have to live above it.

Defender. Don't treat "the instructions" as privileged. Under attention, your system prompt and an attacker's pasted paragraph both bid for the same weight, and the model's learned preference for the system prompt is a tendency, not a boundary you can lean on. Trust boundaries get enforced outside the model (provenance tags, isolation of untrusted content, output filtering) because inside the softmax there are none to enforce.

The KV cache and the tricks that make long context affordable — FlashAttention, sliding-window and sparse patterns, grouped-query attention — get their own module next (Lesson 7), which is also where the compute-versus-memory cost story lands properly. The load-bearing idea to carry out of this one: attention is a learned, soft lookup where every token queries every other, and that "every other" is both the source of its power and the reason it's quadratic.

Sources

  • Vaswani, A. et al. (2017), Attention Is All You Need. arXiv:1706.03762 — scaled dot-product attention (§3.2.1), the √d_k rationale, base-model dimensions (d_model=512, h=8, d_k=d_v=64), and the per-layer complexity comparison in Table 1.
  • NeurIPS 2017 camera-ready of the same paper (papers.neurips.cc/paper/7181-attention-is-all-you-need.pdf). — https://papers.nips.cc/paper/7181-attention-is-all-you-need