Model Families & Architectures · 6 min

Why Decoder-Only Won for Generation

Few-shot prompting turned every NLP task into next-token prediction, retiring per-task fine-tuning and making one decoder cheaper than a zoo of encoders.

Around 2019 the state of the art in NLP looked like a workshop full of specialized jigs. You took a pretrained base and specialized it per task. With an encoder like BERT that usually meant bolting a task-specific prediction head onto the trunk and fine-tuning. With an encoder-decoder like T5 there was no new head to bolt on — the whole point was a text-to-text interface where every task was already framed as string-in, string-out — but you still fine-tuned a separate checkpoint per task. Either way you ran gradient descent on a labeled dataset and shipped a separate set of weights per task. Sentiment got one checkpoint, named-entity recognition got another, translation got a third. It worked, and it produced excellent per-task numbers. It was also operationally miserable. Every new task meant collecting labels, running a training job, versioning yet another artifact, and standing up yet another inference target.

Decoder-only generation didn't win because it was more accurate on any given benchmark. For a while it wasn't. It won because it made that workshop obsolete.

The mental model: one objective, no heads

Start with the training objective, because everything downstream falls out of it. A causal decoder is trained on exactly one task: given tokens t_1 … t_{n-1}, predict a probability distribution over t_n. That's the whole objective. Cross-entropy on next-token prediction, run over a few hundred billion tokens of text.

The trick is that almost any task you care about can be written as text that continues. You don't attach a classifier. You arrange the input so the answer is the natural continuation, then read off the tokens the model generates.

  ENCODER-DECODER / T5 STYLE           DECODER-ONLY STYLE
  (fine-tune a head per task)          (one model, prompt per task)

  base weights                          base weights
     ├── + sentiment head  (train)         │
     ├── + NER head        (train)         │  "Review: ... \n Sentiment:"  --> " positive"
     ├── + QA head         (train)         │  "Translate to French: cat"  --> " chat"
     └── + translate head  (train)         │  "Q: capital of France? A:"  --> " Paris"
                                            │
  N tasks -> N checkpoints              N tasks -> 1 checkpoint, N prompts

T5 (Raffel et al., 2020) already took a big step toward unification. It reframed every task as text-to-text, so the output was always a string rather than a class label or a span. That's the conceptual bridge. But T5's recipe still fine-tuned the model per task: you took the pretrained encoder-decoder and updated its weights on each downstream dataset. The unification was in the format, not yet in the deployment.

The mechanics: in-context learning

GPT-3 (Brown et al., 2020) pushed that format-level unification all the way into inference. Scale the same causal-LM objective to 175B parameters, 96 layers, a 2048-token context, trained on roughly 300B tokens, and something qualitatively new shows up. You can specify the task inside the prompt, give a handful of examples, and the model performs it with no gradient updates at all. They called it in-context learning.

Few-shot means you paste demonstrations straight into the context window:

  Translate English to French:
  sea otter => loutre de mer
  cheese    => fromage
  plush     =>            <- model continues here: " peluche"

Zero gradient steps. Zero new weights. The "learning" happens in the forward pass, conditioned on what's in the context. From an ops perspective this is the whole ballgame. The unit of task specification moved from a training run that produces a checkpoint to a string you write at request time.

Builder: this is why prompt files replaced training pipelines for the long tail of tasks. Adding a capability means editing text and redeploying nothing. The cost of a new task dropped from "label a dataset plus GPU-hours" to "write three examples." That one economic fact is most of why product teams standardized on a single big decoder behind a single endpoint.

The honest numbers

Few-shot GPT-3 was genuinely competitive on some tasks and clearly not on others. The paper is candid about this, and you should be too.

Where it shone: closed-book TriviaQA. Zero-shot GPT-3 scored 64.3%, one-shot 68.0%, few-shot 71.2%. The zero-shot number alone beat a fine-tuned T5-11B by 14.2 points in the same closed-book setting. A model doing no task-specific training beat a model that did. That's the result that reframed the field.

Where it didn't: on SuperGLUE, few-shot GPT-3 was uneven. It roughly matched a fine-tuned BERT-Large on several sub-tasks but trailed the fine-tuned state of the art overall, and it struggled badly on tasks that require comparing two sentences. Word-in-context (WiC) landed near chance, and natural-language inference (RTE, ANLI, CB) stayed weak. If your only metric was peak accuracy on a fixed benchmark with labels available, fine-tuning still won in 2020.

So the win was never "decoder-only is more accurate." It was this: one model gets you most of the way on an enormous range of tasks with no per-task engineering, and you can buy the rest later with fine-tuning when a specific task justifies it. Generality at near-zero marginal task cost beat specialization at high marginal task cost.

Why the architecture, specifically

Two structural reasons the decoder, not the encoder-decoder, became the default carrier for this.

First, clean scaling. A decoder-only model is a single homogeneous stack running one objective on raw text. No paired input/output data, no separate encoder and decoder to balance, no task heads. That uniformity is exactly what the scaling-law work wanted: one loss, one data pipeline, turn the knob. An encoder-decoder splits its parameter budget across two stacks plus a cross-attention bridge to coordinate them — not necessarily more parameters (total budget is chosen independently of the shape), but more moving parts to balance, and for pure generation a lot of that structure is overhead you pay to maintain.

Second, the density of the training signal. Every position in a causal-LM sequence is a supervised prediction, so an n-token document yields n next-token targets. Masked objectives like BERT's only score the roughly 15% of positions they mask. Denser gradient signal per token of data is a real advantage when the whole strategy is "scale the data."

  causal LM:   [the][cat][sat][on] ...   every position predicts the next -> ~n targets
  masked LM:   [the][MASK][sat][MASK]     only masked positions scored     -> ~0.15n targets

Researcher: keep the caveat that "decoder-only won" describes the generation-first, scale-first regime, not a theorem. Encoder-decoders stay strong where input and output are distinct, bounded sequences, so translation and summarization benchmarks still favor T5-family models. Bidirectional encoders remain the right tool for embedding and retrieval, where you want a representation of the whole input rather than a continuation. The field consolidated on decoders because the dominant workload became open-ended generation and instruction-following, not because the other shapes stopped working.

The security angle

The same property that made decoder-only cheap is the root of prompt injection: task specification lives in untrusted text inside the context window. There is no privileged "task head" that a defender controls out-of-band. The instruction and the attacker's data flow through the identical token stream and get identical treatment. A fine-tuned classifier can't be talked out of its task by its input. An in-context model can, because "what task am I doing" is just more tokens.

Defender: treat every token in the context as part of the instruction surface, because architecturally it is. Retrieved documents, tool outputs, and user text are not inert data to a next-token predictor. They are continuation context that steers generation. Enforce trust boundaries around the model (input provenance, output validation, privilege separation on tools) and never assume them inside it.

The deeper point ties back to the objective. A causal decoder only ever does one thing: predict the next token conditioned on everything before it. Few-shot learning, instruction following, and prompt injection are the same mechanism seen from three angles. That collapse of "many tasks" into "one conditional distribution" is exactly what the next lesson on encoder versus decoder attention masks makes precise. The causal mask is the small architectural choice that makes all of this hold together.

Sources