Infrastructure, Hardware & Production Deployment · 6 min
Measuring the Right Latency
Tokens-per-second alone is a vanity metric. TTFT, inter-token latency, and cost per successful task are what a user and a budget actually feel.
A model that streams 200 tokens per second can still feel broken. If the first token takes four seconds to arrive, the user has already tabbed away. Tokens per second is the number vendors put on slides because it's big and easy to measure. It's also the number that hides the two things a person actually experiences: how long they stare at a blank box, and how smoothly text arrives once it starts. Those are different clocks, driven by different parts of the machine, and you have to measure them separately or you're measuring nothing.
Two clocks, not one
An LLM request has two distinct phases, and each has its own latency signature.
Prefill processes your prompt. The model reads every input token at once, in parallel, building the KV cache. This is compute-bound and it scales with prompt length. A 50-token question prefills almost instantly; a 30,000-token document dump does not. Prefill ends the moment the first output token is ready.
Decode generates the answer one token at a time. Each new token depends on the last, so this phase is serial and memory-bandwidth-bound: the GPU spends most of its time reading weights and the growing KV cache out of HBM, not doing math. Decode speed sets how fast text flows after it starts.
That split gives you the two metrics that matter for anything interactive:
- TTFT (Time To First Token) — from request sent to first token received. Prefill-dominated, plus any queuing delay before your request even starts. This is the blank-box wait.
- ITL (Inter-Token Latency), sometimes called TPOT (Time Per Output Token) — the average gap between consecutive output tokens during decode. This is the smoothness of the stream.
request ──▶ [queue] ──▶ [prefill] ──▶ first token ──▶ [decode...] ──▶ done
└──────────── TTFT ────────────┘ └─ ITL per token ─┘Here's why the aggregate lies. Suppose two servers each finish a 500-token answer in the same total time. Server A: TTFT 300 ms, then a steady 15 ms between tokens. Server B: TTFT 3,000 ms, then a blistering 9 ms between tokens. Both report a similar end-to-end "tokens/sec." Server A feels responsive and readable. Server B feels frozen, then dumps a wall of text faster than anyone reads. Same throughput number, opposite product.
There's a human-facing ceiling hiding in ITL. Comfortable reading runs roughly 4–7 words per second, and a word is about 1.3 tokens, so sustaining around 10 tokens/sec of perceived output is plenty for a chat UI. Past that, faster decode buys you nothing a reader notices; it only matters for batch jobs or agents consuming the output programmatically. Knowing that ceiling stops you from over-optimizing decode while TTFT rots.
Throughput is a server metric, latency is a user metric
Tokens/sec and requests/sec describe the fleet, not the person. They're the right lens for capacity planning and cost, the wrong lens for UX. The two are in direct tension, and the knob between them is concurrency: how many requests the server batches together at once.
Modern servers like vLLM and TensorRT-LLM use continuous (in-flight) batching. They merge many requests into one running batch and add or drop sequences token by token. Push more concurrent requests into a batch and aggregate throughput climbs — the GPU stays busy. But every request in that batch now shares the same decode step, so each individual user's ITL gets worse, and new arrivals wait longer in the queue, so TTFT gets worse. You are trading per-user latency for fleet efficiency.
This is why a single number is a trap. Report throughput at a stated latency bound: "3,000 tokens/sec while keeping p95 TTFT under 500 ms." A throughput figure with no latency constraint is just a server running at a concurrency that makes users miserable.
concurrency ↑ → tokens/sec ↑ (good for cost)
TTFT ↑, ITL ↑ (bad for user)
the operating point is a choice, not a factThe average is where UX goes to die
Report a mean latency and you have described a server almost no one experiences. LLM latency distributions have long right tails. A request that lands mid-batch during a prefill spike, or one that triggers a KV-cache eviction, can take many times the median. Averages smear those events into invisibility.
Use percentiles. p50 is the typical request. p95 and p99 are the tail: the slowest 5% and 1%.
p50 TTFT: 240 ms ← what a demo shows you p95 TTFT: 900 ms ← what a busy afternoon shows you p99 TTFT: 4200 ms ← the users writing angry tickets
Tail latency isn't an edge case you can wave off, because users don't send one request. A conversation is ten turns; an agent chain is fifty tool calls. If each step has a 1% chance of hitting your p99, a 50-step agent run hits it with probability 1 - 0.99^50 ≈ 0.40. The tail is the median experience of any multi-step workload. MLPerf Inference codifies this: its server-scenario LLM benchmarks are only valid if the system meets TTFT and TPOT constraints at a required percentile. You don't get to report throughput unless the tail behaved.
Measure percentiles per phase — TTFT and ITL separately — and measure them under realistic load, not on an idle box. A p99 TTFT taken with one request in flight tells you nothing about production.
Utilization, errors, and the only cost that counts
Three more signals turn a latency chart into an operational picture.
GPU and memory utilization. "GPU util" from nvidia-smi is misleading: it reports whether a kernel was running, not whether the hardware was busy doing useful work. A decode-bound server can show 95% util while the compute units idle, waiting on HBM bandwidth. The number that actually governs capacity is KV-cache memory pressure. Each concurrent sequence holds a KV cache proportional to its context length; when that pool fills, the server stops admitting requests (queue grows, TTFT spikes) or preempts and recomputes sequences. Watch KV-cache occupancy and request queue depth, not just util percent.
Error rate. Timeouts, out-of-memory kills, truncated streams, malformed tool calls, content-filter rejections. A server that's fast because it drops 3% of requests under load is not fast. Error rate has to sit next to your latency percentiles, or the percentiles are computed on survivors only — a classic way to make a bad system look great.
Cost per successful task. This is the metric that ties it all together, and it's almost never cost per token. A user doesn't want tokens; they want a correct summary, a working diff, a resolved ticket. So the denominator is completed, acceptable outcomes, and the numerator includes everything you actually paid for:
cost per successful task =
(GPU-hours × $/hour, including failed & retried attempts and wasted prefill)
─────────────────────────────────────────────────────────────────────────────
number of tasks that actually succeededA cheaper-per-token model that needs three retries, longer prompts, or a bigger context to get the answer right can cost more per success than a pricier one that nails it first try. Cost per token flatters the wrong model. Cost per successful task is what your budget feels at the end of the month, and it's the number to put in front of anyone choosing between two setups.
Put it together and the scorecard for an interactive LLM service reads: TTFT p95, ITL p95, throughput at that latency bound, error rate, KV-cache occupancy, and cost per successful task. Tokens/sec can stay on the slide. Just don't mistake it for the truth.
Sources
- Reddi, V. J., et al. "MLPerf Inference Benchmark." 47th Annual International Symposium on Computer Architecture (ISCA), 2020. arXiv:1911.02549. https://arxiv.org/abs/1911.02549
- MLCommons. "MLPerf Inference Rules" (server-scenario latency constraints, including TTFT and TPOT). GitHub: mlcommons/inference_policies. https://github.com/mlcommons/inference_policies/blob/master/inference_rules.adoc
- Kwon, W., et al. "Efficient Memory Management for Large Language Model Serving with PagedAttention." ACM Symposium on Operating Systems Principles (SOSP), 2023. arXiv:2309.06180. https://arxiv.org/abs/2309.06180
- vLLM Documentation. "Production Metrics." https://docs.vllm.ai/en/latest/serving/metrics.html
- Databricks. "LLM Inference Performance Engineering: Best Practices." 2023. https://www.databricks.com/blog/llm-inference-performance-engineering-best-practices
- Nielsen, J. "Response Times: The 3 Important Limits." Nielsen Norman Group, 1993. https://www.nngroup.com/articles/response-times-3-important-limits/