Emerging Frontiers · 7 min

Two Ways to Scale: Training-Time Versus Inference-Time Compute

For a decade, "make it better" meant "make it bigger." The newer axis spends compute at request time instead — and it changes the cost curve completely.

GPT-2 shipped in 2019 with 1.5 billion parameters. A year later GPT-3 had 175 billion — a hundredfold jump in twelve months, and a training run estimated in the millions of dollars. The bet behind that leap was simple and, for a while, unbeatable: pour in more parameters, more data, and more training FLOPs, and the model gets better in a way you can predict on a log-log plot. That bet defined the frontier for most of a decade. It is no longer the only game, and this module is about the other one.

The old axis: bigger is a plan, not a hope

The reason scale worked so well is that it was measurable. Kaplan and colleagues at OpenAI showed in 2020 that test loss falls as a smooth power law in three quantities — model size N, dataset size D, and training compute C — across seven orders of magnitude. You could draw a straight line on log axes and extrapolate. That turned "make it better" from an art into a budgeting exercise: pick a compute budget, read off the loss you should expect, provision the GPUs.

The 2020 paper's advice was parameter-heavy: given more compute, spend most of it making the model bigger and undertrain it slightly. Two years later, Hoffmann and colleagues at DeepMind revisited the exact tradeoff and found the earlier recipe left performance on the table. Their result — the "Chinchilla" scaling laws — is worth stating concretely, because it changed how labs spend money.

Kaplan (2020):     given 10x compute, scale params ~5.5x, data ~1.8x
Chinchilla (2022): given 10x compute, scale params ~3.2x, data ~3.2x
                   -> params and tokens should grow roughly equally

Chinchilla itself was the proof by example: 70B parameters trained on 1.4 trillion tokens beat Gopher (280B parameters) across a broad benchmark suite, at the same training budget. A model four times smaller won because it was fed the right amount of data. The lesson everyone took away — data and parameters should scale in roughly equal measure — is why the models you use today are trained on trillions of tokens rather than hundreds of billions.

But notice what both papers optimize. They optimize a number you pay for once, at training time, and then amortize across every future request. A Chinchilla-optimal model is cheap to serve precisely because it is small for its quality. The entire framing assumes the interesting compute is spent before the model ever meets a user.

The new axis: spend at request time

Here is the reframe. When you ask a model a hard question, the default behavior is to emit tokens left to right, one forward pass per token, and stop. The amount of computation is fixed by the length of the answer, not the difficulty of the question. A model spends the same effort per token deciding 2 + 2 as it does on a competition math problem. That is obviously wrong, and inference-time scaling is the family of techniques that fixes it: give the model more computation to solve a particular request, and let it use that computation to think.

The simplest version is repeated sampling. Instead of one answer, draw k answers and pick the best.

def best_of_n(model, prompt, n, verify):
    candidates = [model.sample(prompt, temperature=0.8) for _ in range(n)]
    # verify() scores each: a unit test, a reward model, or a majority vote
    return max(candidates, key=verify)

When you can check answers cheaply — code that either passes tests or doesn't, math with a known answer — this is startlingly effective. Brown and colleagues (2024) found that on coding and math benchmarks, the fraction of problems solved by at least one of k samples keeps climbing as k grows into the hundreds, often as a clean power law in the number of samples. On one coding benchmark, coverage rose from a small fraction of problems at k=1 to the large majority at k in the hundreds — same model, frozen weights, no retraining. Just more tries. The catch hides in the word check: running k=100 samples costs literally a hundred times the inference FLOPs of a single answer, so the trick only pays for itself when a verifier can pull the one good answer out of the pile. With no verifier, the fallback is a majority vote across samples, which helps far less.

The subtler version is longer reasoning. Chain-of-thought prompting (Wei et al., 2022) showed that asking a model to write out intermediate steps improves accuracy on multi-step problems, because each generated token becomes scratch space the next token can attend to. Modern reasoning models — OpenAI's o1, then a wave of open counterparts — take this further by being trained to produce long internal reasoning traces before answering, and to spend more of those tokens when a problem is hard. The knob you turn at inference is "how long should it think," and reported accuracy on math and coding rises with it.

Why this reframes the frontier

The reason people are excited is not that either technique is magical in isolation. It is that they trade against each other, and the trade is favorable in places the old axis was stuck.

Snell and colleagues (2024) put numbers on it. They ask a pointed question: given a fixed pool of FLOPs, are you better off spending it to pretrain a bigger model, or to run a smaller model harder at inference? Their answer is "it depends, and the dependence is predictable." On problems where a smaller base model is already in the right ballpark, optimally allocated test-time compute can beat a model up to 14 times larger. On the hardest problems, or when you need low latency, pretraining still wins. The practical output is a policy: for a given prompt and difficulty, there is a compute-optimal split between thinking longer and being bigger.

That is the whole shape of the module in one idea. Two dials:

             cheap to serve                 expensive to serve
train-time  |------------------------------------------------->
            (pay once, amortize forever)

             fast, cheap answer             slow, thorough answer
infer-time  |------------------------------------------------->
            (pay per request, scales with difficulty)

One caution before you treat any of this as gospel: these are empirical regularities fit to the models and benchmarks we have, not laws of physics. Chinchilla revised Kaplan within two years, and the exact test-time curves shift with the task, the base model, and how good your verifier is. The shape of the tradeoff holds up; the precise numbers are not promises.

The economics are genuinely different, and you should feel the difference before the rest of the module gets into mechanism. Training compute is a capital expense: enormous, one-time, and paid by whoever trained the model. Inference compute is a per-request operating expense you pay every single time, and it shows up as latency the user waits through and dollars on your serving bill. A model that "thinks" for thirty seconds to nail a proof is wonderful for offline research and unacceptable for autocomplete. Best-of-1000 sampling is a bargain when a verifier is free and a catastrophe when every sample costs a cent and nothing can check the answer.

So the questions this module keeps returning to are the ones this reframe forces:

  • Where does the compute go? More parameters, more thinking tokens, more parallel samples, or an external tool or memory that does the work instead.
  • Can you verify? Repeated sampling only pays off when something — tests, a reward model, a majority vote — can tell good from bad. Without a verifier, more samples is just more noise.
  • What is the latency budget? Inference-time methods buy quality with wall-clock time. Fine for an agent working overnight; fatal for a chat cursor.

Keep those three in hand. State-space models change the cost of the forward pass itself. Test-time training and memory change what "solving a request" even includes. Self-improving systems try to feed inference-time work back into the weights. All of them are moves on the same board: the frontier is no longer a single dial labeled "bigger," and knowing which dial a new result actually turns is how you read the research without getting fooled.

Sources