Emerging Frontiers · 7 min

Inference-Time Reasoning Systems

o1 and R1 did not get smarter by getting bigger — they spend more compute per question. The durable idea is extra inference-time computation and intermediate reasoning state: sampling, verification, and search, not whether the UI shows the "thinking."

OpenAI previewed o1 in September 2024, and the jump it made over GPT-4o did not come from training a bigger network. OpenAI never disclosed o1's size, but its own framing was blunt: the model learned to think before it answers. What changed is what happens after you hit send. o1 spends seconds to minutes generating and revising a long internal reasoning trace before it commits to an answer. On the 2024 AIME math competition, GPT-4o scored about 13%. o1 scored 83%. Nobody trained a 10x larger network to get there. They spent more compute per question.

That is the whole idea of inference-time reasoning, and it is worth separating from the marketing around it. The durable concept is this: a request can buy extra computation and carry intermediate reasoning state, and quality can rise with that spend even when the weights are frozen. Everything else — the visible "thinking" text, the branded reasoning modes — is presentation layered on top of that mechanism.

Spend at inference, not just at training

For most of the deep-learning era, the recipe for a smarter model was a bigger model trained on more data. That is training-time scaling: expensive and one-shot. You pay once, ship the weights, and every query then costs the same fixed forward pass.

Inference-time scaling moves a second dial. Hold the weights fixed and vary how much compute you spend answering a single question. Snell et al. (2024) studied this directly and found that for a fixed compute budget, letting a smaller model "think longer" at inference can beat a much larger model answering in one pass — in their FLOPs-matched tests, a model up to 14x smaller came out ahead once the test-time compute was allocated well. The catch is in their own results: the win depends on problem difficulty. Easy questions gain little from extra sampling; hard ones gain a lot. There is no regime where more thinking always helps.

Concretely, "spend more compute" takes a few shapes:

1. Sample many candidate answers, then pick one.
2. Generate a long intermediate reasoning trace before answering.
3. Search over reasoning steps, keeping promising branches, pruning dead ones.
4. Generate a draft, critique it, and revise (self-correction).

These are not exclusive. A production reasoning system usually stacks several.

Sampling and verification: the simplest lever

The cheapest way to trade compute for quality is to sample the same question k times and combine the results. Two classic combiners:

# self-consistency: majority vote over sampled reasoning traces
answers = [model.sample(prompt, temperature=0.8) for _ in range(k)]
final = most_common(extract_answer(a) for a in answers)

# best-of-n with a verifier: score each candidate, keep the best
cands = [model.sample(prompt) for _ in range(n)]
final = max(cands, key=verifier.score)

Self-consistency (Wang et al., 2022) just votes. Best-of-n needs a verifier — a separate model or checker that scores candidates. The verifier is the load-bearing part. If you can cheaply tell whether an answer is good, sampling many and filtering is enormously effective. If you cannot, sampling just hands you many confident wrong answers to choose badly among.

This is why math and code drove the early reasoning wins. A math answer can be checked against a known result; code can be run against tests. These are verifiable outcomes — the correctness signal is mechanical, not a matter of taste. For open-ended work like "write a good strategy memo," there is no unit test, and the same techniques degrade because the verifier is now itself a fallible model scoring another fallible model.

How o1 and R1 learned to do it

Sampling at inference is something you bolt on from outside. The bigger step was training models to produce good long reasoning traces on their own, so that a single sample is already a structured, self-correcting chain rather than a lucky guess.

OpenAI's o1 was trained with reinforcement learning to think before it answers, producing a long internal chain of thought that it learns to refine, backtrack within, and error-check. OpenAI has been explicit that it hides the raw chain of thought from users and shows only a summary — which is the cleanest possible proof that the reasoning and the displayed text are different things.

DeepSeek-R1 (DeepSeek-AI, 2025) is the version you can read the recipe for, because they published it. Their first system, R1-Zero, was trained by pure RL on a base model with no supervised reasoning examples at all — just a reward for producing verifiably correct answers in the required format. The reward was rule-based:

reward = accuracy_reward   # did the final answer match the known solution?
       + format_reward     # was the reasoning wrapped in <think>...</think>?

No neural reward model scored correctness. A learned reward model is something the policy can hack, and for math and code you do not need one — you can just check the answer. Under this pressure the model spontaneously grew its reasoning: over training it began generating longer traces, and the paper describes an "aha moment" where R1-Zero learns to stop, re-examine its approach, and try again. Nobody hand-wrote that behavior. It emerged because longer, self-correcting reasoning got rewarded. The numbers back the story: over the course of RL training, R1-Zero's pass@1 on AIME 2024 climbed from 15.6% to 71.0%, and reached 86.7% with majority voting over samples — with no supervised reasoning data anywhere in that loop.

R1-Zero's raw output was hard to read — mixed languages, messy formatting — so the shipped DeepSeek-R1 added a small cold-start dataset and a later formatting stage. But the core claim holds and is reproducible: give a capable base model a verifiable reward and enough RL, and it will learn to spend inference compute well.

The part everyone gets wrong: chain of thought is not reasoning

Here is the trap. Because o1 and R1 surface a stream of "thinking…" text, people conflate three separate things:

  1. The visible chain-of-thought text in the UI.
  2. The intermediate reasoning state the model actually computes over.
  3. Whether the model does extra inference-time computation at all.

Only (3) is the durable idea. The visible text is a rendering choice. OpenAI hides o1's real trace and shows a summary; DeepSeek shows the trace; a model could do all of this internally and show you nothing. None of that changes how much compute was spent or how good the answer is.

Two practical consequences follow. First, do not read the chain of thought as a faithful log of why the model answered. It is text sampled from the same model — often revealing, but capable of rationalizing an answer the model reached by other means. Anthropic and others have documented reasoning traces that leave out the actual cause of an answer. Second, do not judge "is this a reasoning system?" by whether a thinking panel appears. A model that silently samples 32 candidates and verifies them is doing more inference-time reasoning than one that prints a tidy paragraph and answers in a single pass.

What this costs you

None of this is free, and the bill comes due at request time.

Single-pass answer:     ~1x tokens,   ~1x latency
Reasoning trace:        5-50x tokens, seconds to minutes
Best-of-32 + verifier:  ~33x compute for one answer

You pay for reasoning tokens — o1 and R1 both bill hidden thinking tokens as output — and you pay in latency, where a reasoning model can take 30 seconds or more against a chat model's one. So the engineering choice is a routing decision, not an "always on" one:

  • Route hard, verifiable, high-value queries to the reasoning path. Competition math, tricky debugging, multi-step planning where a checker exists.
  • Keep easy or latency-sensitive queries on the single-pass model. Snell's own data shows extra compute barely helps easy problems; you would be paying 30x for nothing.
  • Invest in the verifier. Wherever you can build a real checker — tests, a validator, a ground-truth key — do it. That is what makes sampling and search pay off.

The frontier here is genuinely unsettled. How far verifiable-reward RL generalizes beyond math and code, whether these traces are faithful, and where the compute-versus-quality curve flattens are all open questions as of 2025. Treat vendor claims about "reasoning" as claims about inference-time compute and verification, and ask which one they actually mean.

Sources