Data, Pre-Training & Post-Training · 6 min

Scaling Laws: Why Chinchilla Says Data Is the Binding Constraint

How compute-optimal scaling works, why Kaplan 2020 pushed the field toward oversized undertrained models, and why Chinchilla reframed data as the real bottleneck.

A scaling law is an empirical claim with teeth. Hold the recipe fixed, vary one input across orders of magnitude, and the model's test loss falls along a clean power law. That says far more than "bigger is better." A power law lets you fit a curve on the models you can afford to train and extrapolate to the one you can't yet afford. That extrapolation is how anyone justifies spending eight figures on a single training run before it has produced a single token.

Two inputs do the real work: N, the parameter count, and D, the number of training tokens. Both cost compute, and the standard approximation for a dense transformer is

C  ≈  6 · N · D      FLOPs

The 6 is bookkeeping: roughly 2 FLOPs per parameter on the forward pass and 4 on the backward pass, summed over every token. That one equation turns a vague question, "how should I scale?", into a constrained optimization. Given a fixed compute budget C, choose the (N, D) that minimizes loss. Everything below is an argument about where that minimum sits.

Kaplan 2020: loss bends to power laws

Kaplan and colleagues at OpenAI varied each input in isolation and measured two power laws:

L(N) ≈ (N_c / N)^0.076        L(D) ≈ (D_c / D)^0.095

The exponents are small, so returns arrive slowly. A 10x increase in parameters buys only about a 10^0.076 ≈ 1.19x cut in the reducible loss. Kaplan's headline was not either exponent on its own. It came from tracing the most efficient path through the compute-loss frontier, which told them to grow parameters much faster than data:

N_opt ∝ C^0.73        D_opt ∝ C^0.27

Read literally: as compute grows, pour about 73% of the growth into parameters and only 27% into tokens. Their own advice was to "train very large models on a relatively modest amount of data and stop well before convergence." The field took it to heart. GPT-3 (175B) saw roughly 300B tokens. Gopher (280B) saw roughly 300B. Megatron-Turing NLG reached 530B parameters on a similar diet. The parameter count was the trophy, and the token count barely moved.

Chinchilla 2022: the correction

Hoffmann and colleagues at DeepMind re-ran the study more carefully, over 400 models, with one fix that turned out to matter: the learning-rate schedule was matched to each run's token count. Kaplan had reused a single schedule, which quietly penalizes the runs trained on more tokens. Instead of two isolated curves, they fit a joint surface:

L(N, D)  =  E  +  A / N^α  +  B / D^β
E = 1.69    A = 406.4    B = 410.7    α = 0.34    β = 0.28

E is the irreducible loss, the entropy floor of natural language that no amount of training removes. The other two terms are the parameter-limited and data-limited penalties. This time α and β come out close to equal, and that flips the answer. Minimize the surface under C = 6ND and the budget splits evenly:

N_opt ∝ C^0.50        D_opt ∝ C^0.50

Parameters and tokens should grow together. Put real numbers in and you get the rule everyone quotes: about 20 training tokens per parameter at the compute-optimal point.

Fixed compute budget  C = 6·N·D  is a hyperbola in (N, D):

  D (tokens)
    ^
1.4T|  *  <- Chinchilla-optimal (70B, 1.4T)  ~20 tok/param
    |    \
    |      \
300B|        *  <- Gopher-style (280B, 300B)  ~1 tok/param
    |          \____
    +--------------------------> N (params)
        70B      280B

Same compute C. Same curve. Very different loss.
Gopher sat in the wrong corner: too many params, starved of data.

Then came the payoff. DeepMind trained Chinchilla, 70B parameters on 1.4T tokens, using the same compute as Gopher's 280B. The smaller model won. It beat Gopher, GPT-3, Jurassic-1, and the 530B MT-NLG across the board, reaching 67.5% on MMLU against Gopher's roughly 60%, a jump of more than seven points. Four times fewer parameters, a better model, identical training cost. The parameters in those larger models were not wrong so much as underfed.

Here is the reframe worth carrying. For a compute-optimal run, data is the binding constraint. Parameters are a checkbook problem, because you can always buy more. Tokens are finite. There is only so much non-garbage text in existence, and once you fix N the optimizer wants about 20N of it. That scarcity is why the race after Chinchilla stopped being "who has the biggest model" and became "who has the most, and cleanest, tokens."

Researcher: Don't cite "20 tokens/param" as a law of nature. The 20x ratio is specific to the tokenizer, data distribution, and architecture in that paper. A 2024 replication by Besiroglu, Erdil, Barnett and You (Epoch AI) reproduced the core result but showed that Chinchilla's third estimation method reported confidence intervals so tight they would have required hundreds of thousands of runs, when the authors likely trained fewer than 500. Refitting brought that method back in line with the other two. Treat 20x as a strong prior, not a constant.

What actually goes wrong

People quote the ratio and stop reading. Chinchilla-optimal minimizes training compute for a target loss and says nothing about inference. If you plan to serve a model to millions of requests, the smart move is to overtrain a smaller model well past 20x, paying extra training cost once to cut inference cost forever. That is why Llama-style models push to extreme ratios. An 8B trained on about 15T tokens sits near 1,900 tokens per parameter, roughly 90x past "optimal." That is not a mistake, it is a different objective. Sardana and colleagues (2023) formalize it: fold expected inference volume into the budget and the optimum shifts hard toward smaller-and-longer.

Extrapolation is a promise, not a guarantee. The curve gets fit inside one compute range and assumed to hold an order or two beyond it. Data quality moves the constants, repeated-epoch training breaks the "D fresh tokens" assumption, and the irreducible term E means the curve eventually flattens. Chasing the last sliver of loss then costs exponentially more compute for a benchmark bump nobody feels.

"Big model" tells you almost nothing on its own. A parameter count without a token count is uninterpretable. A 70B trained on 200B tokens and a 70B trained on 2T are different animals wearing the same size label.

Builder: Before you fine-tune or self-host, find the token count, not just the parameter count. A capable "small" model trained long is cheaper to run and often matches a starved larger one, for the same reason a right-sized model beats an oversized one at fixed compute. When you count "your" tokens for continued pretraining, count after tokenization; the same corpus yields a different D under a different tokenizer (see the tokenization lesson in this module).
Defender: Chinchilla's real security consequence sits upstream. If data is the binding constraint, every lab scrapes the entire reachable web to hit its token target, which drops untrusted, attacker-writable text into the training set by construction. The data-poisoning and backdoor surface grows precisely because tokens are scarce and curation can't keep up. The leverage in an LLM's threat model has shifted from the architecture to the corpus, and provenance plus filtering (the data-curation lesson) are now front-line controls rather than hygiene.

The clean version to keep in your head: loss follows power laws in both N and D, a fixed compute budget is the hyperbola C ≈ 6ND, and the minimum-loss point on that curve wants N and D scaled together, near 20 tokens per parameter. Kaplan's frontier analysis aimed the field at giant, undertrained models. Chinchilla's corrected fit showed most of them were parked in the wrong corner of the same curve, and that the scarce resource was never parameters. It was data.

Sources