Model Families & Architectures · 6 min
Three Architectures: Encoder-Only, Decoder-Only, Encoder-Decoder
The transformer split into three families along two axes, attention masking and training objective, and those two choices explain what each model can and can't do.
Every transformer you'll meet is built from the same block: self-attention plus a feed-forward network, stacked N times. What makes BERT unlike GPT-3 isn't the block. It's two decisions layered on top: which tokens each position may look at (the attention mask), and what the model is trained to predict (the objective). Get those two axes straight and you can place almost any checkpoint on sight, including ones that didn't exist when you read this.
The two axes
Attention masking governs how information flows inside a layer. Under bidirectional attention, token 5 can attend to tokens 1 through 10, the whole sequence, past and future alike. Under causal (autoregressive) attention, token 5 sees only tokens 1 through 5. Everything to its right is masked with -inf before the softmax, so future tokens contribute exactly zero. This is a structural property baked into the forward pass, not a tuning knob.
Bidirectional (BERT) Causal (GPT-3)
attends to attends to
1 2 3 4 5 1 2 3 4 5
1 ■ ■ ■ ■ ■ 1 ■ · · · ·
2 ■ ■ ■ ■ ■ 2 ■ ■ · · ·
3 ■ ■ ■ ■ ■ 3 ■ ■ ■ · ·
4 ■ ■ ■ ■ ■ 4 ■ ■ ■ ■ ·
5 ■ ■ ■ ■ ■ 5 ■ ■ ■ ■ ■
(see everything) (lower triangle only)Training objective decides what signal the weights actually learn from. Three canonical choices map cleanly onto the three families.
Encoder-only: BERT and masked language modeling
BERT (Devlin et al., 2019) stacks bidirectional encoder blocks: 12 layers and 110M parameters for BERT-base, 24 layers and 340M for BERT-large. Its objective is masked language modeling (MLM). Corrupt 15% of the input tokens, then predict the originals from the context on both sides.
The 15% split hides a detail people forget. Of the chosen tokens, 80% become [MASK], 10% become a random token, and 10% are left untouched.
Input: the cat sat on the mat Corrupt: the [MASK] sat on the mat Target: cat
Why bother with the random and unchanged tokens? Because [MASK] never appears at inference time. A model that only ever saw [MASK] would learn to produce output only where the mask token sits, then fall apart on clean text. Perturbing some of the kept-real tokens forces every position to build a useful representation, not just the masked slots.
Bidirectionality is exactly why BERT is a representation model rather than a generation model. Each output vector is conditioned on the entire sentence, which is ideal for classification, named-entity recognition, retrieval embeddings, and reranking. What you can't do is sample text left to right: position 5 was trained assuming it could see position 6, so there's no coherent next-token distribution to roll forward.
Builder: When the task is score, classify, or embed this text, an encoder like BERT (or a modern descendant like DeBERTa or an e5 embedding model) is smaller, faster, and often more accurate than prompting a decoder LLM. A 110M-parameter encoder on a CPU beats a 7B decoder at sentiment for a fraction of the cost.
Decoder-only: GPT-3 and next-token prediction
GPT-3 (Brown et al., 2020) goes the other way: a stack of causal decoder blocks, 96 layers and 175B parameters, trained on next-token prediction. The objective could not be simpler. Given tokens 1 through k, predict token k+1. The causal mask keeps it honest, because the model literally cannot cheat by peeking at the token it's meant to guess.
Input: the cat sat on the
Predict: cat sat on the mat
(every position predicts its own next token, in parallel, during training)That one objective is why decoder-only models generate: prediction and use are the same operation. Feed a prefix, sample token k+1, append it, repeat. And because every position yields a training signal, unlike MLM where only ~15% of positions do, the objective is dense and scales beautifully. That density is a big part of why the field consolidated here.
The GPT-3 paper's real headline wasn't the architecture but in-context learning: drop a few worked examples into the prompt and the model performs the task with no gradient updates. Everything now called prompting is a downstream consequence of that finding.
Defender: In-context learning has no privilege boundary. The same mechanism that lets a prompt demonstrate a task lets injected text redefine it. Instructions and data ride the same token stream into the same causal context. That's the architectural root of prompt injection, not a bug to be patched at the model level. (The security module covers this in depth.)
Encoder-decoder: T5 and span corruption
T5 (Raffel et al., 2020) keeps both halves. A bidirectional encoder reads the input; a causal decoder writes the output, and cross-attention lets each decoder step look back at the full encoded input. The baseline runs ~220M parameters, roughly a BERT-base-sized encoder bolted to a matching decoder.
Its framing is the whole point: text-to-text. Every task, whether translation, summarization, or classification, is cast as string in, string out. Pre-training uses span corruption: drop contiguous spans (15% of tokens, mean length ~3) and replace each with a single unique sentinel token. The decoder then emits the missing spans, each tagged by its sentinel.
Input: Thank you <X> me to your party <Y> week Target: <X> for inviting <Y> last <Z>
Here <X> and <Y> are sentinels and <Z> marks the end. This runs leaner than BERT's token-by-token masking: one sentinel can stand in for several dropped tokens, so the sequences stay short.
The encoder-decoder split earns its keep when input and output are genuinely different objects, like a French sentence and its English translation, or an article and its summary. The bidirectional encoder gets to digest the whole source before a single output token is committed. The cost is roughly double the parameters and a two-stage forward pass, which is much of why decoder-only won the scaling race. For open-ended generation, one causal stack is simpler and cheaper to serve.
The placement table
FAMILY MASK OBJECTIVE GENERATES? ARCHETYPE encoder-only bidirectional masked-LM no BERT, DeBERTa, e5 decoder-only causal next-token yes GPT-3, Llama, Mistral encoder-decoder bi + causal span corruption yes T5, FLAN-T5, BART*
*BART corrupts differently (denoising) but shares the encoder-decoder shape.
Researcher: The tidy objective-implies-capability story has been dented on purpose. UL2 and PrefixLM mix objectives inside one model, and encoder-only checkpoints have been retrofitted for generation. The two axes are still the right first question, just not the last word. When a model surprises you, re-check which mask and which objective it actually trained under.
Next time you open a model card, hunt for exactly two facts: causal or bidirectional attention, and what the model was trained to predict. Almost everything else, whether it can embed, whether it can generate, whether it will hallucinate a continuation of your classification prompt, falls out of those two answers.
This pairs with the tokenization lesson (what a token even is before any masking happens) and sets up the scaling laws lesson, which explains why the industry poured its compute into the causal-decoder corner of this table.
Sources
- Devlin, Chang, Lee, Toutanova (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL. https://aclanthology.org/N19-1423/ (arXiv:1810.04805)
- Raffel et al. (2020). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5). JMLR. https://arxiv.org/abs/1910.10683
- Brown et al. (2020). Language Models are Few-Shot Learners (GPT-3). NeurIPS. https://arxiv.org/abs/2005.14165
- Vaswani et al. (2017). Attention Is All You Need. https://arxiv.org/abs/1706.03762
- Lewis et al. (2020). BART: Denoising Sequence-to-Sequence Pre-training. https://arxiv.org/abs/1910.13461