Data, Pre-Training & Post-Training · 6 min

From Base Model to Assistant: Supervised Fine-Tuning

How a small set of hand-written demonstration pairs and the ordinary next-token loss turn a text-completing base model into one that follows instructions.

A base model is a very good autocomplete. Feed GPT-3 the string Explain how DNS works and it will not answer you. It predicts the most probable continuation of that string as the string tended to appear across its training corpus. That continuation might be another exam question, a table of contents, or a Stack Overflow header. The knowledge is in there. The behavior of answering a request is not. Supervised fine-tuning (SFT, also called instruction tuning) is the cheap, blunt, load-bearing step that installs the behavior.

The mental model: reshaping, not teaching

Hold onto one idea and most of SFT falls out of it. You are not adding knowledge. You are shifting a conditional distribution. The base model already assigns some probability to a helpful answer following your question. It just assigns more probability to other continuations. SFT takes a small pile of (prompt, ideal response) examples and nudges the weights so that, once the input is wrapped in a chat format, the helpful continuation becomes the likely one.

That is why SFT works with so little data. LIMA (Zhou et al., 2023) aligned a 65B LLaMa with 1,000 curated examples, no reinforcement learning of any kind, and produced responses that matched or beat much heavier pipelines in human comparisons. The authors call the underlying idea the Superficial Alignment Hypothesis: essentially all capability is learned in pre-training, and alignment mostly teaches the model which sub-distribution to speak from and what format to speak in. Treat that as a strong claim to test rather than settled fact, but it explains the shape of the field well.

The mechanics: same loss, different data

Here is the part that surprises people. SFT uses the exact same objective as pre-training: token-level cross-entropy, predict the next token given every previous one. No new loss function. No new architecture. The whole transformation comes from what you compute that loss over, which is curated demonstrations wrapped in a consistent chat template.

Two implementation details do the real work.

Formatting with role delimiters. Every example is serialized with special tokens marking who is speaking:

<|user|>
Explain how DNS resolves a hostname.
<|assistant|>
DNS resolution walks a hierarchy...
<|end|>

The delimiters are ordinary tokens the model learns to condition on. <|assistant|> becomes the cue that means "a helpful answer starts now." <|end|> teaches the model to stop. Without a learned stop token your assistant runs on forever and starts answering its own follow-up questions.

Loss masking on the prompt. You compute gradients only on the response tokens and mask out the prompt:

tokens:  <|user|> Explain DNS <|assistant|> DNS  walks  the  ... <|end|>
mask:      0    0    0    0        0        1    1     1    1     1
                (no gradient here)         (gradient flows here)

Skip the mask and the model spends capacity learning to generate plausible user questions, which is the opposite of what you want. I have watched exactly that quietly degrade an otherwise clean run.

InstructGPT: SFT as step 1

The canonical reference is Ouyang et al., 2022, Training language models to follow instructions with human feedback, the InstructGPT paper and the recipe behind ChatGPT's first generation. RLHF gets the headlines, but step 1 is plain SFT, and it does a large share of the visible work.

The honest numbers:

  • About 40 contractors, hired through Upwork and Scale AI and screened for agreement with the researchers' judgments, wrote demonstrations by hand. This is expensive artisanal data, not a scrape.
  • The SFT set was roughly 13,000 prompts. The reward-model stage used about 33,000 and PPO about 31,000, so SFT is the smallest of the three.
  • Prompts came in three flavors. Plain prompts were arbitrary tasks labelers invented. Few-shot prompts paired a task with several worked examples. User-based prompts were drawn from real requests submitted to the OpenAI API waitlist. That last bucket matters: training on the distribution of things people actually ask beats training on the tasks you imagined they would ask.
  • Training ran 16 epochs with cosine learning-rate decay and residual dropout of 0.2.

That epoch count deserves a pause. The SFT models overfit validation loss after a single epoch, and OpenAI kept training to 16 anyway, because human-preference ratings and reward-model score kept improving well past the point where the loss curve said stop. Final checkpoints were chosen by reward-model score on held-out data, not by validation loss. The lesson generalizes: for SFT, the loss curve is a weak proxy for the thing you actually care about. Judge on behavior.

   base model                SFT (step 1)              [later: RM + PPO]
 next-token over    ──►   16 epochs on 13k     ──►   preference tuning
 the whole internet      hand-written demos          (next lesson)
   "completes text"        "follows format"           "ranked better"

Why format and quality beat quantity

Three things go wrong in practice, and none of them are fixed by adding more data.

Inconsistent formatting poisons the signal. If half your demonstrations end responses with <|end|> and half just trail off, the model learns a fuzzy, probabilistic stop and will sometimes ramble. The template has to be byte-for-byte identical between training and serving. A mismatched system-prompt wrapper at inference time is the single most common reason a freshly fine-tuned model appears to "ignore everything it learned."

Style is contagious, and that includes the bad parts. Demonstrations always answer. Train on that and the model learns that the correct move is to produce a confident answer even when the honest response is "I don't know." At the data level you are teaching a measured dose of hallucination and sycophancy. If you want the model to hedge or refuse, those behaviors have to appear in the demonstrations. A model cannot infer a behavior it never saw modeled.

A few bad examples do outsized damage. With only 1k to 13k examples and 16 epochs, every example is seen many times over. One demonstration carrying a subtly wrong fact or a lazy format gets memorized, not averaged away.

Builder: Before you scale the dataset, read 50 of your own examples end to end and grade them like a hostile reviewer. A 500-example set you personally audited beats a 50k set you scraped. That is the whole LIMA result in one sentence.
Defender: SFT data is an injection surface. Whoever writes or filters demonstrations can plant a backdoor ("when the prompt contains token X, emit Y") that survives into production and stays nearly invisible in eval. Treat demonstration provenance the way you treat dependency provenance: signed, reviewed, access-controlled. Fine-tuning on user-submitted "corrections" without review is a data-poisoning vector wearing the costume of a feature.
Researcher: The open question is what SFT actually moves. Probing studies suggest it mostly re-weights features the base model already has rather than writing new ones, which fits LIMA. How much genuine new capability SFT can add, versus only surfacing what is latent, is unsettled. Measure it on your own tasks instead of assuming an answer.

A worked micro-run

Say you want a base model to emit structured incident summaries. You write 300 examples, each an (alert log → JSON summary) pair in one fixed template. You mask the prompt, train three to five epochs, and watch a held-out set for whether the JSON parses and the fields are right, not the loss number. That is a complete, real SFT loop. The base model already "knew" what a severity level was. You taught it to emit one on cue, in your shape.

SFT gets you a model that follows instructions. It does not get you a model that reliably picks the better of two acceptable answers, or one that holds a refusal when a user pushes back, because a single demonstration can only show one right answer and never a ranking between two. Closing that gap is exactly the job of preference-based post-training (RLHF and DPO), and it is the next lesson.

Sources

From Base Model to Assistant: Supervised Fine-Tuning — All About LLMs, from AI to Z · AdversariaLLM