Start Here — LLMs in Plain Language · 5 min

How a model learns: pretraining, then made helpful

A model is built in two stages — first it reads a mountain of text to predict what comes next, then it's taught to be helpful — and knowing that explains why it can be brilliant and wrong in the same breath.

Nobody sat down and typed the facts into a language model. Nobody wrote a rule that says "if the user asks for a recipe, respond politely." That surprises people, because the thing acts like it was carefully programmed. It wasn't. Its behavior grew out of a training process, in two distinct stages. Understanding those two stages is the single best way to understand why the model is so capable and so unreliable at the same time.

If you take one idea from this lesson, take this: a model is not taught facts, it is trained to predict text — and everything it seems to "know" is a side effect of getting good at that one game.

Stage one: pretraining, the prediction game

The first stage is called pretraining — the long, expensive phase where a raw model reads an enormous amount of text and learns to continue it.

Here's the whole game. Show the model a stretch of text with the next piece hidden, and make it guess that next piece. Then reveal the answer. If the guess was off, nudge the model's internal settings a tiny bit so it would have guessed a little better. Now do that again. And again. Billions of times, across a huge slice of the written internet, plus books, code, and more.

Those "internal settings" are called parameters — numbers, billions of them, that together decide how the model turns input into a guess. Think of them as billions of tiny dials. Training is nothing more than turning those dials, a hair at a time, until the model is genuinely good at predicting what comes next.

Why does a guessing game produce something that seems to know things? Because to predict the next word well, you have to pick up on the patterns underneath the words. To finish the sentence "The capital of France is ___," it helps to have absorbed that Paris and France go together. To continue a paragraph of working code, it helps to have absorbed how the code is structured. The model was never told any of this directly. It backed into a rough, statistical version of "knowledge" because that knowledge made it better at the prediction game. That's the whole engine.

This is also why sheer size isn't the only thing that matters. Researchers found that to get the most out of a model, you have to grow the amount of training text alongside the size of the model — feed a big model too little text and you've wasted it. Capability comes from the model and the data together, not from size alone.

And here's the catch that trips everyone up: the model only ever saw text up to a certain date. That's the training cutoff. Pretraining happens once, on a frozen snapshot of text, and then it's over. The model has no memory of anything after that snapshot — not last week's news, not this morning's score. When it answers a question about a recent event, it is either guessing from stale patterns or repeating something a tool fetched for it in the moment. The knowledge is baked in at training time and does not refresh on its own.

What a raw pretrained model actually does

At the end of stage one you don't have a helpful assistant. You have a very good text-continuer, and that's a stranger thing than it sounds.

Ask a raw pretrained model "What is the capital of France?" and it might not answer you. It might reply with "What is the capital of Germany? What is the capital of Spain?" — because on the internet, that question often shows up in a list of similar questions, so continuing with more questions is a perfectly reasonable prediction. The model isn't being difficult. It's doing exactly its one job: continue the text plausibly. It has no idea it's supposed to be a helpful assistant, because nobody has told it that yet.

Stage two: making it helpful

The second stage turns that raw text-continuer into something you'd actually want to talk to. It has a couple of parts.

First, fine-tuning on examples (often called instruction tuning). Fine-tuning just means a further round of training on the model you already have. Here, it's shown many examples of the pattern "here's a request, here's a good response," so it learns the shape of being an assistant: when you get an instruction, you follow it, instead of rattling off more instructions.

Second, learning from human preferences — the part usually labeled RLHF, which just means "reinforcement learning from human feedback." In plain terms: the model produces a few different answers, people (or a system trained to imitate people's judgments) mark which answer is better, and the model is nudged toward the kind of answers people prefer. Do this a lot and the model drifts toward being helpful, clear, and — importantly — toward clearly refusing requests that are obviously harmful. This is where the "personality" and the manners come from.

Notice what's still true even here: nobody wrote if-then rules. The refusals and the helpfulness were shaped by feedback, the same way the knowledge was shaped by prediction. It's all learned behavior, all the way down.

Why this explains the weirdness

Once you see the two stages, the model's strangest habit makes sense. It can hand you a fluent, confident, completely wrong answer — a hallucination — and sound exactly as sure as when it's right.

That's not a bug that slipped past someone. It falls straight out of how the thing was built. The model was trained to produce plausible text, not true text. Plausible and true usually travel together, which is why it's right so often. But when they part ways, the model has no separate sense of truth to fall back on. It will confidently continue with whatever pattern fits best, correct or not.

So the same machinery gives you both the brilliance and the confident mistakes. They aren't two different systems. They're the same guessing game, seen from two sides.

That guessing game runs on text the model has been chopped up in a particular way. Next you'll see the very first thing that happens when you send a message — how your words get broken into the pieces the model actually reads.

How a model learns: pretraining, then made helpful — All About LLMs, from AI to Z · AdversariaLLM