Data, Pre-Training & Post-Training · 7 min
What Pre-Training Actually Optimizes: Next-Token Prediction
Pre-training minimizes one loss, cross-entropy on the next token, and that single objective explains perplexity, base models, and why raw prediction yields broad capability.
In Module 2 you built the machine: a transformer that reads a sequence of token embeddings and, at each position, emits a vector of logits over the whole vocabulary. This lesson is about the one number that machine is tuned to minimize. There is no zoo of objectives and no auxiliary "understand the text" loss. Pre-training optimizes a single quantity, the cross-entropy of the next token, over trillions of tokens. Everything you experience as capability falls out of that.
The objective, concretely
Take a document and tokenize it into a sequence t₀, t₁, t₂, …, tₙ. The model reads a prefix and, at every position, predicts a probability distribution over what comes next. The training signal is almost embarrassingly simple. At position i, how much probability did the model put on the token that actually came next, tᵢ₊₁?
This is self-supervised, so nobody labels anything. The text is both the input and the answer key. One chunk of length L+1 yields L prediction problems, all scored in parallel thanks to the causal attention mask from Module 2 (each position may attend only to positions at or before it, so the model can't cheat by reading ahead).
sequence: The cat sat on the
input: The cat sat on <- positions the model reads
target: cat sat on the <- the "next token" at each position
(labels are inputs shifted left by one)An autoregressive model factors the probability of the whole sequence into a product of conditionals:
P(t₀…tₙ) = Π P(tᵢ | t₀ … tᵢ₋₁)
We want that product to be large for real text, meaning the model should find genuine documents likely. Maximizing the product is maximizing likelihood. Take the log so products become sums and underflow goes away, negate it so we can minimize, average per token, and you have the loss:
loss = − (1/N) Σ log P(tᵢ | t<ᵢ)
That is cross-entropy between the model's predicted distribution and a one-hot target that puts all its mass on the observed token. Maximum likelihood and minimum cross-entropy are the same knob seen from two sides.
A worked micro-example
Say the true next token is sat. The model's softmax assigns it probability 0.5, so the per-token loss is −ln(0.5) ≈ 0.693 nats. Assign 0.01 instead and the loss jumps to −ln(0.01) ≈ 4.605. Assign 1.0, perfect and correct, and the loss is 0. The penalty grows without bound as confidence in the right answer approaches zero, which is exactly the pressure you want. Catastrophic overconfidence in the wrong token gets punished hard.
What does a loss of 0.693 feel like? Exponentiate it: exp(0.693) = 2. That is perplexity, the exponential of the average cross-entropy, and its reading is clean. It's the effective number of equally-likely choices the model is deciding between at each step. Perplexity 2 means the model is as uncertain as a fair coin flip. Perplexity 20 means it has narrowed a roughly 50,000-token vocabulary down to about 20 live candidates per position.
Two anchors set the scale. A model that has learned nothing spreads mass uniformly over a ~50k vocabulary, giving perplexity ≈ 50,000 and loss ≈ 10.8 nats. A strong modern model on ordinary English lands in the low tens or single digits, depending on the corpus. That whole journey, 50,000 down to about 10, is what the loss curve traces during training, and it stays smooth precisely because the objective never changes underneath it.
Builder: watch the units. Frameworks compute cross-entropy in nats (natural log); much of the literature reports bits (log base 2). perplexity = exp(loss_in_nats) = 2^(loss_in_bits). Mixing them silently is a classic way to convince yourself your model is twice as good, or half as good, as it really is.Why one dumb loss induces broad capability
Here is the part that surprises people. "Predict the next token" sounds like autocomplete, yet it produces translation, arithmetic, code, and chains of reasoning. The reason is that predicting the next token well across a huge, diverse corpus eventually makes memorizing surface statistics stop paying off. To finish "The capital of France is ___" you need a fact. To finish "import numpy as ___" you need code conventions. To finish "2380 + 4176 = ___" you need to have internalized addition, because the corpus holds far too many distinct sums to store. To finish a murder mystery's last line you need to have tracked the plot. Compression is the pressure. The model has finite parameters and cannot keep the corpus verbatim, so the only way to drive loss down is to learn the generative structure that produced the text. Prediction, pushed hard enough, forces understanding-shaped machinery to appear as a side effect.
This is also why scaling is lawful rather than magical. The Chinchilla team fit loss as a function of parameters N and training tokens D:
L(N, D) = A / N^α + B / D^β + E
Loss falls predictably as you add parameters or data, sinking toward an irreducible floor E, the entropy of language itself: the part no model can predict because it is genuinely a coin toss over which of several valid words a writer happened to pick. Their headline practical result was that for a fixed compute budget you should spend it on roughly 20 training tokens per parameter, not on a giant model trained briefly. Compute runs about C ≈ 6ND.
Researcher: the smooth L(N,D) surface is what lets small ablations extrapolate. Change the data mix or architecture at 1B params, read the loss delta, and you can often predict the sign, and roughly the size, of the effect at 70B before spending the money.What a "base model" actually is
Run pre-training to convergence and you get a base model, a pure next-token predictor. It is not an assistant. Hand it "What is the capital of France?" and it might answer "Paris", or it might continue with "What is the capital of Germany? What is the capital of Spain?", because in its corpus a question is very often followed by more questions (quiz sheets, worksheets, exam banks). It is completing the document faithfully, with no notion that you wanted an answer.
raw corpus ──► [ pre-training: minimize next-token cross-entropy ] ──► BASE MODEL
│ completes text
[ SFT on demonstrations + RLHF ] ◄────────────────┘ (Module 3, later)
│
▼
ASSISTANT (what users chat with)Almost nobody outside a lab talks to a base model. The chat model you use has been through supervised fine-tuning on instruction/response demonstrations and then preference optimization (RLHF and its relatives). InstructGPT is the canonical published example, fine-tuned on about 13k human-written demonstrations plus a reward model trained over ranked completions. Those later stages shift the distribution toward helpfulness. They surface and shape what pre-training already installed far more than they teach new facts. That is why the base model's loss is a ceiling: post-training can make a model cooperative and safe, but it cannot conjure knowledge that predicting the corpus never put there.
Defender: the objective is the attack surface. A base model has no concept of "instruction" versus "content" versus "system prompt." It only continues text. Every prompt injection and jailbreak ultimately exploits that. The alignment layer is learned behavior sitting on a substrate that just wants to complete the sequence, and adversarial input tries to make continuing the injected text the most likely thing to do. The loss is why these defenses are statistical rather than absolute.
Carry this into the next lesson: perplexity and the loss curve are not abstractions bolted onto the transformer. They are the transformer's own logit outputs scored against reality, one token at a time. When we turn to how corpora get built and cleaned, keep it in view. Everything the model can ever do is bounded by what minimizing this one loss over that specific text could teach it. Garbage in the corpus is not a metaphor here; it is literally the distribution the model is being optimized to reproduce.
Sources
- Kaplan et al., "Scaling Laws for Neural Language Models" (2020), arXiv:2001.08361
- Hoffmann et al., "Training Compute-Optimal Large Language Models" (Chinchilla, 2022), arXiv:2203.15556
- Ouyang et al., "Training Language Models to Follow Instructions with Human Feedback" (InstructGPT, 2022), arXiv:2203.02155 / NeurIPS
- Sebastian Raschka, "How does next-token prediction train a large language model?" — sebastianraschka.com — https://sebastianraschka.com/faq/docs/next-token-prediction.html
- Chip Huyen, "Evaluation Metrics for Language Modeling," The Gradient (2019) — thegradient.pub — https://thegradient.pub/understanding-evaluation-metrics-for-language-models/
- Abby Morgan, "Perplexity for LLM Evaluation," Comet (2024) — comet.com — https://www.comet.com/site/blog/perplexity-for-llm-evaluation/
- Epoch AI, "Chinchilla Scaling: A Replication Attempt" (2024) — epoch.ai — https://epoch.ai/blog/chinchilla-scaling-a-replication-attempt