Foundational Mechanics · 6 min

The Transformer block, one mental model

A transformer is one running per-token vector that every block reads from and adds back to; hold that picture and the rest follows.

If you remember one thing about how a transformer is wired, make it this. For each token position there is a single running vector, and every part of the network reads from it and writes back to it by addition. That is the whole communication model. Attention heads, feed-forward layers, the embedding at the bottom, the unembedding at the top: none of them talk to each other directly. They leave messages on a shared bus. Elhage and colleagues at Anthropic named that bus the residual stream, and it is the most useful picture I know for reasoning about these networks without drowning in matrices.

The stream, and what rides on it

Follow one token's journey. It enters as an embedding, a vector of d_model numbers, the very same per-token vector the previous module (L2) built from the token's meaning and position. Vaswani's 2017 model used a width of 512; a modern 7B model runs closer to 4096, and frontier models go wider still. That embedding is not a second, separate object from the stream: it is the stream's initial value. From here up we call it the residual stream, it flows straight up through the whole network, and no block ever replaces it. Each block computes a small update and adds it in.

   logits  <- unembedding reads the final stream
     ^
   [ + ]<---- MLP block N writes an update
     |
   [ + ]<---- Attention block N writes an update
     |
    ...        the stream is never overwritten,
     |         only added to, layer after layer
   [ + ]<---- MLP block 1
     |
   [ + ]<---- Attention block 1
     ^
  embedding  (token + position) seeds the stream

A block has two moves. Attention lets a position pull information from other positions. It is the only operation in the architecture that mixes across the sequence. The MLP, a two-layer feed-forward network usually about 4x wider in the middle, processes each position on its own with no cross-talk. Attention decides where to look; the MLP decides what to make of it. Each attention head learns a lower-dimensional view of the residual stream. The MLP does almost the opposite: it usually expands each token into a wider intermediate space, applies its nonlinear transformation, then projects the result back to the residual-stream width. Either way the block reads from the stream and adds its result back in.

That "read a subspace, write a subspace" pattern is the payload of the residual-stream framing. The stream might have 4096 dimensions, but no single block uses all of them for one purpose. A head reads a learned lower-dimensional projection of the stream, not a fixed physical slice; an MLP neuron writes along one particular direction. Different features live in different directions, and because everything is additive, you can trace an end-to-end path (embedding, to a specific head, to the logits) as a plain sum of contributions. That decomposability is exactly why interpretability researchers can isolate circuits at all.

Why the "+" matters, and why the norm

The addition is load-bearing, not cosmetic. Vaswani writes each sub-layer as LayerNorm(x + Sublayer(x)): the input x is carried forward untouched and the sub-layer's output rides on top. That skip connection is what makes deep stacks trainable. During backprop the gradient gets a clean path straight down the stream, because the derivative of x + f(x) with respect to x always contains a 1. Strip the skip out and gradients threading dozens of matrix multiplies either vanish or explode, and the network never learns. Residual connections are the reason we can stack 32 or 80 blocks instead of six.

Normalization is the other half of the stability story. Before a block reads the stream, its input gets re-centered and re-scaled to a consistent magnitude. Think of a level control on the bus: however loud the stream has grown by layer 40, the block sees a signal in a predictable range. Two practical wrinkles hide behind the original formula.

  • Placement moved. Vaswani put the norm after the addition (post-norm). Nearly every modern LLM puts it before the sub-layer (pre-norm), so the residual path stays a clean, un-normalized highway and only the copy the block reads gets normalized. Pre-norm trains far more stably at depth, which is why 100-layer models are feasible at all.
  • The norm got simpler. Many current models use RMSNorm, which scales without mean-centering. It is cheaper and works about as well.

Builder: load a Llama-style checkpoint and you will find input_layernorm and post_attention_layernorm per layer. That is a pre-norm block with one norm before attention and one before the MLP. The names describe position in the block, not what gets normalized, and misreading them is a classic reimplementation bug.

What actually goes wrong

The model earns its keep when things break.

Activations that grow. Because the stream is a running sum, its magnitude tends to climb with depth, since every layer adds energy. On long or unusual inputs, values can drift into ranges that overflow in fp16, and that is a genuine cause of NaN logits mid-generation. The fix is numerical (bf16, or a careful norm epsilon), but the reason is structural: an additive stream accumulates.

A few directions carry the load. Empirically, a handful of stream dimensions and the attention on the first token carry enormous magnitude, where the model parks attention it does not need. Quantize naively, crush those outliers, and quality falls off a cliff. The stream is not uniformly important; some directions are structural.

Defender: the residual stream is also where you intervene. Because contributions add, you can inject a steering vector at some layer and shift behavior. That is the mechanism behind activation steering and behind probes that ask whether a model represents that it is being tested. If your threat model includes model tampering, the injection surface is a single additive site, not a tangle of weights.

A tiny lab you can run

You do not need to train anything to see the stream. Load a small open model with a library that exposes hidden states (output_hidden_states=True on most HF models) and feed it a sentence. You get back one vector per layer per token, which are snapshots of the stream. Do two things with them.

First, take the layer-by-layer stream for the final token, multiply each snapshot by the unembedding, and read the top token at every depth. You will watch the prediction form: early layers make syntactic guesses, later layers lock onto meaning. This is the logit lens.

Second, compute the norm of the update each block adds, meaning the difference between consecutive hidden states. Attention and MLP contributions vary wildly by layer. Some blocks barely write anything; others dominate. That variance is the network showing you where its work happens.

Researcher: the logit lens works only because of the residual stream. The final unembedding reads the stream linearly, so applying it to an intermediate stream is a legitimate, if lossy, readout. In an architecture without a shared additive stream, the trick would be meaningless.

The one-sentence version

A transformer is a stack of blocks, each adding a small attention update and a small MLP update to one shared per-token vector, with norms to keep the signal sane and skip connections to keep the gradients alive. Hold that picture, and attention itself (the QK "where to look" and the OV "what to copy") is just the next zoom level in, which the following module opens up.

Sources

The Transformer block, one mental model — All About LLMs, from AI to Z · AdversariaLLM