Foundational Mechanics · 7 min

Decoding: how the next token is chosen

The decoder collapses a vocabulary-sized score vector into one token: softmax, temperature, top-k/top-p, and why temperature 0 still isn't deterministic.

A language model never picks a word. At each step it produces a vector of scores, one per token in its vocabulary, and there are tens of thousands of them. Something downstream collapses that vector into a single concrete token. That something is the decoder. It is the part of the stack almost nobody inspects, yet it controls most of what you experience as the model's personality: crisp or rambling, safe or surprising, repeatable or not.

Nothing in this lesson touches the weights. We are only deciding how to read them — one position at a time, because generation is autoregressive: the token the decoder picks is appended to the running input and fed back in, and the model runs again for the next position. That loop is what Lesson 7 (KV caching) exists to make fast; here we only care about the single decision made at each step of it.

From logits to a distribution

The raw score vector is called the logits. They are unbounded real numbers, some positive, some negative, and on their own they mean little. Softmax makes them comparable by exponentiating and normalizing, so every value is positive and the whole set sums to 1:

p_i = exp(logit_i / T) / Σ_j exp(logit_j / T)

Set T to 1 for now. Softmax is the bridge from "scores" to "a probability distribution over the vocabulary." Four candidate tokens, worked out:

token     logit    exp(logit)   prob
" the"     2.0      7.39         0.63
" a"       1.0      2.72         0.23
" cat"     0.5      1.65         0.14
" xylo"   -3.0      0.05         0.004
                    ----------
                    sum ≈ 11.81

Now there is a real distribution to work with. The decoder either samples a token from it or refuses to sample and takes the top one. Every named strategy is a variation on that single choice.

Greedy, and why "take the best" isn't free

Greedy decoding takes the argmax, the highest-probability token, at every step. Given identical arithmetic it is deterministic, and it is often what you want for extraction, classification, or code, where one answer is correct.

Greedy is also myopic. The locally most-probable token at each step does not add up to the most-probable sequence, and for open-ended text it produces something worse than either: bland output that loops. Holtzman et al. (2019) showed that maximization-based decoding, greedy and beam search alike, drives models into degenerate repetition ("I don't know. I don't know. I don't know."). Their sharper point: human text does not ride the high-probability ridge the model wants to walk. Real language keeps dipping into lower-probability tokens. Always chasing the maximum yields prose that is statistically likely and unmistakably robotic.

Temperature reshapes before you sample

Back to T. Temperature scales the logits before softmax. It adds no randomness of its own; it changes the shape of the distribution that randomness later draws from.

  • T < 1 sharpens the distribution. The top token dominates.
  • T > 1 flattens it. The tail gets a real vote.
  • T → 0 collapses everything onto the peak, which is exactly greedy.

Same four tokens at three temperatures:

             T=0.5     T=1.0     T=2.0
" the"       0.84      0.63      0.46
" a"         0.11      0.23      0.28
" cat"       0.04      0.14      0.22
" xylo"     ~0.00      0.004     0.04

At T=0.5 the model says " the" almost every time. At T=2.0 the nonsense token " xylo" climbs from a 0.4% chance to roughly 4%, a tenfold jump, and " cat" rises to nearly match " a". This is the first knob people reach for and the first they misuse. Cranking it up for "creativity" mostly hands probability mass to the unreliable tail.

Top-k and top-p truncate the tail

That tail is the problem temperature cannot fix, because temperature reshapes globally. It has no way to say "these ten tokens are plausible and the other 49,990 are garbage." Truncation can.

Top-k keeps the k highest-probability tokens, zeros the rest, renormalizes, and samples. GPT-2 shipped with k=40. Simple, but rigid: k is fixed while the model's confidence is not. Sometimes only two tokens make sense and k=40 drags in 38 bad ones; sometimes fifty are all fine and k=40 amputates good options.

Top-p (nucleus) sampling is the Holtzman et al. contribution, and it fixes that rigidity. Instead of a fixed count, keep the smallest set of tokens whose probabilities sum to at least p, then sample from that set. The candidate pool breathes with the model's confidence.

sorted probs:  0.63  0.23  0.14  0.004  ...
cumulative:    0.63  0.86  1.00
p = 0.9  ->  nucleus = { the, a, cat }   (0.86 < 0.9, add next; stop)
p = 0.7  ->  nucleus = { the, a }        (0.63 < 0.7, add; 0.86 ≥ 0.7, stop)

When the model is certain, the nucleus is one or two tokens and you get near-greedy behavior for free. When it is genuinely uncertain, the nucleus widens and diversity returns, without ever inviting the flat garbage tail that raw temperature would. For chat, top_p around 0.9 to 0.95 with a modest temperature is the workhorse. The two compose: temperature reshapes, top-p then clips.

Builder. Order of operations matters and providers differ. The usual pipeline is temperature, then top-k, then top-p, then sample. Set both temperature and top_p and you are stacking two effects; move one at a time or you will never know which knob did what. For anything you parse (JSON, function calls, labels) use temperature 0 and stop tuning.

Why one prompt gives different answers

Two independent causes. Conflate them and you will burn an afternoon.

Sampling. With any T > 0 and a non-trivial nucleus you are drawing a random token. A different draw gives a different continuation, and because the autoregressive loop feeds each token back in as the condition for the next, a tiny divergence at one step compounds into an entirely different paragraph. That is by design.

The seed. Sampling runs on a pseudo-random generator. Fix its seed (many APIs expose a seed parameter) and you reproduce the exact draw. In principle.

What temperature 0 does and does not buy you

Temperature 0 removes the sampling cause. No sampling, pure argmax. People assume that buys determinism. It does not, and this catches serious teams.

The argmax is only as stable as the logits underneath it, and those logits are not bit-stable between runs. The dominant cause, pinned down by Thinking Machines' 2025 analysis, is not "GPU floating-point is random." It is batch-size dependence. Your request gets batched with whatever other traffic is live; GPU reduction kernels tile their sums differently for different batch sizes; and floating-point addition is non-associative, so (a+b)+c and a+(b+c) can differ in the last bits. Usually the argmax survives. But when the top two tokens are nearly tied, a last-bit wobble flips the winner, and every token after that point diverges.

So temperature 0 guarantees you take the top token. It does not guarantee the top token is the same one next run, and it says nothing about reproducibility across model versions, hardware, or batching regimes. Batch-invariant kernels, released alongside that analysis and integrated with vLLM, can deliver bit-identical output, but that is a deliberate, non-default engineering choice.

Defender. "Deterministic at temp 0" is a shaky foundation for a safety control. A refusal classifier or guardrail that passes in testing can flip on a near-tie under production batching. Test with realistic concurrency, and if you need a hard reproducibility guarantee for audit logs or eval gates, pin the seed and run batch-invariant kernels. Do not trust temp 0 alone.
Researcher. Report the whole decoding config: temperature, top_k, top_p, seed, engine, and version. A benchmark delta of one or two points can be pure decoder variance. Fu et al. (2026) show token-probability nondeterminism is largest exactly in the 0.1–0.9 probability band, and Atil et al. (2024) measured accuracy swinging by up to 15% across nominally deterministic runs, so even greedy runs are not a clean control group.

Try it

Take any local model with logit access. Feed it one prompt and dump the top-10 logits with their softmax probabilities. Recompute the probabilities by hand at T = 0.7, 1.0, 1.5 and watch the tail swell. Then run the same prompt 20 times at T=0, batched against filler traffic, and diff the outputs. They will probably agree, right up until the one prompt sitting near a logit tie, which won't. That single disagreement is the whole lesson.

Decoding is where a distribution becomes a decision. Later modules on reproducible evals and prompt-injection defenses both lean on it: you cannot reason about a model's output distribution until you control how it gets collapsed.

Sources