Infrastructure, Hardware & Production Deployment · 7 min
Prefill Versus Decode: Two Phases, Two Bottlenecks
Processing the prompt and generating tokens are different workloads — one compute-bound, one memory-bandwidth-bound. Optimize them as if they were one and you lose both.
Send a 2,000-token prompt to a model and watch the terminal. There's a pause where nothing happens, then text streams out word by word at a steady clip. Those two behaviors — the pause and the stream — aren't one system running at two speeds. They're two different computations with two different bottlenecks, and the reason your GPU is bored during one and starved during the other comes down to a single ratio.
The pause and the stream
An autoregressive transformer generates text in two phases. The first is prefill: the model ingests your entire prompt at once. Every one of those 2,000 tokens gets embedded, run through all the attention and MLP layers, and its key and value vectors get computed and stashed in the KV cache. Because all the prompt tokens are known up front, this happens in parallel. The attention and projections for the whole prompt collapse into a handful of large matrix multiplications.
The second phase is decode. Now the model produces output one token at a time. It takes the last token, runs it through every layer, samples the next token, appends it, and repeats. Each decode step processes exactly one new token — but to compute attention for that token it must read every key and value already sitting in the cache, which by now holds the whole prompt plus everything generated so far.
That structural difference is the whole lesson. Prefill is a wide, parallel, compute-heavy burst. Decode is a long, skinny, sequential dribble where each step does very little arithmetic but touches a lot of memory.
Here's the shape in pseudocode:
# PREFILL: one shot, all prompt tokens together
hidden = embed(prompt_tokens) # shape [seq_len, d_model]
for layer in model.layers:
k, v = layer.attn.kv(hidden) # big matmul, [seq_len, d_model]
kv_cache[layer].store(k, v) # fill the cache
hidden = layer(hidden) # attention over seq_len tokens
# DECODE: loop, one token at a time
token = sample(hidden[-1])
while token != EOS:
h = embed(token) # shape [1, d_model]
for layer in model.layers:
k, v = layer.attn.kv(h) # tiny matmul, [1, d_model]
kv_cache[layer].append(k, v) # grows by one row
h = layer.attn(h, kv_cache[layer]) # read the WHOLE cache
token = sample(h)Look at the matmul sizes. In prefill the projections operate on a [seq_len, d_model] matrix. In decode the same projections operate on [1, d_model] — a single row. The weight matrices are identical and enormous. What changes is how much work you extract per byte of weight you load.
Arithmetic intensity is the whole game
The concept that separates the two phases is arithmetic intensity: FLOPs performed per byte moved from memory. A GPU has a peak compute rate (roughly 990 TFLOP/s of BF16 on an H100 SXM) and a peak memory bandwidth (about 3.35 TB/s of HBM3). Divide them and you get the machine's balance point — near 300 FLOPs per byte. Run above that intensity and you're compute-bound. Run below it and you're memory-bandwidth-bound, and your expensive tensor cores sit idle waiting for data.
Prefill lives above the line. When you multiply a [2000, d_model] activation matrix by a weight matrix, you load each weight once and reuse it across all 2,000 rows. High reuse, high intensity, compute-bound. The GPU runs hot.
Decode lives far below the line. Each step multiplies a [1, d_model] vector against the same giant weight matrices. You stream billions of parameters out of HBM to produce exactly one token, then reload them from scratch for the next step. Reuse is essentially one. A matrix-vector product does two FLOPs — one multiply, one add — per parameter, and in FP16 each parameter is two bytes, so arithmetic intensity sits around 1 FLOP per byte: a couple hundred times under the balance point. Decode is memory-bandwidth-bound, full stop.
This is why the roofline model is the right mental picture. Prefill sits up on the compute ceiling; decode sits down on the bandwidth slope. Same model, same GPU, opposite constraints.
A back-of-envelope check makes it concrete. A 7B model in FP16 is about 14 GB of weights. To generate one token in decode you must stream all 14 GB through the ALUs at least once. On a 3.35 TB/s GPU that's a hard floor near 14 / 3350 ≈ 4.2 milliseconds per token from weight loading alone, before you count a single KV-cache read. That floor is bandwidth, not compute. Buy a GPU with twice the FLOPs and the same bandwidth and it barely moves.
TTFT versus inter-token latency
Two numbers dominate serving, and each is owned by one phase.
Time to first token (TTFT) is how long you wait for the stream to start. That's the pause: prefill, plus scheduling overhead. TTFT scales with prompt length because prefill does work proportional to the number of prompt tokens — and the attention term inside it grows quadratically with sequence length, so very long prompts hurt more than linearly. A 200-token question feels instant; paste a 32,000-token document in front of it and the same model can sit silent for seconds before the first token appears. If your users paste long inputs, prefill is your TTFT problem.
Inter-token latency (ITL), sometimes called time-per-output-token, is the gap between streamed tokens. That's decode. It's set by memory bandwidth and by how big the KV cache has grown, because every step re-reads it. ITL barely cares how many FLOPs your prompt cost; it cares how big the model is and how much cache each step must scan.
TTFT -- prefill -- compute-bound -- grows with prompt length ITL -- decode -- bandwidth-bound -- grows with model size + KV cache
Confusing the two is the classic mistake. A team complains that generation is slow, throws more compute at it, and nothing improves — because the bottleneck was decode bandwidth, and FLOPs were never the constraint.
Why this splits your optimization strategy
Once you accept that prefill and decode are different workloads, the standard serving tricks stop looking like a grab bag and start looking like targeted answers.
Batching helps decode enormously. Decode wastes bandwidth by loading the whole model to serve one token. Load those same weights once and serve 64 requests' tokens in the same step, and you've raised arithmetic intensity roughly 64x at almost no extra weight traffic. That's why throughput-oriented serving batches aggressively, and why continuous batching — a scheduler slotting new requests into a running batch, as vLLM does — is such a large win: it keeps the decode batch full. Prefill, already compute-bound, gains far less from batching and can even hurt latency by monopolizing the GPU. The asymmetry is stark: a batch of one leaves decode running at a few percent of peak FLOPs, while a full batch can lift aggregate throughput by an order of magnitude for the same per-step weight read.
Prefill and decode fight each other on shared hardware. A long prefill is a compute hog; drop it into the same batch as ongoing decodes and it stalls everyone's stream, spiking their ITL. Two research directions attack this. Chunked prefill breaks a big prompt into pieces and interleaves them with decode steps so no single prefill starves the decoders; Sarathi-Serve is the canonical implementation. Disaggregation goes further and runs prefill on one pool of GPUs and decode on another. Splitwise and DistServe both showed that because the two phases have opposite resource profiles, you can pick different hardware and different parallelism for each and hit better goodput under latency targets than any co-located configuration.
The KV cache is the decode-phase resource to guard. It grows linearly with total sequence length, and every decode step reads it, so it eats both memory capacity and the bandwidth that already bottlenecks you. PagedAttention, the mechanism behind vLLM, manages it in non-contiguous blocks to kill fragmentation and pack in more concurrent sequences. Multi-query and grouped-query attention shrink it outright by sharing key/value heads across query heads. These are all decode optimizations; none of them touch prefill's compute problem.
The practical takeaway: measure TTFT and ITL separately, always. If TTFT is your pain, look at prompt length, prefill parallelism, and chunking. If ITL is your pain, look at model size, KV-cache footprint, batch size, and bandwidth. Tune them as if they were one workload and you'll reach for the wrong knob — and lose both.
Sources
- Samuel Williams, Andrew Waterman, David Patterson. "Roofline: An Insightful Visual Performance Model for Multicore Architectures." Communications of the ACM, 2009. https://dl.acm.org/doi/10.1145/1498765.1498785
- Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al. "Efficient Memory Management for Large Language Model Serving with PagedAttention." SOSP 2023. https://arxiv.org/abs/2309.06180
- Pratyush Patel, Esha Choukse, Chaojie Zhang, et al. "Splitwise: Efficient Generative LLM Inference Using Phase Splitting." 2023. https://arxiv.org/abs/2311.18677
- Yinmin Zhong, Shengyu Liu, Junda Chen, et al. "DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving." OSDI 2024. https://arxiv.org/abs/2401.09670
- Amey Agrawal, Nitin Kedia, Ashish Panwar, et al. "Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve." OSDI 2024. https://arxiv.org/abs/2403.02310
- Noam Shazeer. "Fast Transformer Decoding: One Write-Head is All You Need." 2019. https://arxiv.org/abs/1911.02150
- vLLM Documentation. https://docs.vllm.ai/en/latest/