Infrastructure, Hardware & Production Deployment · 7 min
The Inference Optimization Toolkit
Quantization, speculative decoding, and the parallelism strategies: what each one buys you, what it costs, and when it is the wrong tool. Diagnose the bottleneck first, then pick the one technique that moves it.
A 70B model in FP16 wants about 140 GB just to hold its weights, before a single token of KV cache. That number is why this toolkit exists. Every technique below is a different answer to the same pressure: you have a model, a GPU (or eight), a latency target, and a bill, and they refuse to all fit at once. There is no "make it fast" button. There is a bottleneck, and each tool moves a different one, sometimes at the expense of another. The skill is diagnosis before treatment.
So start by naming the bottleneck. Are you out of memory (the model won't load, or the KV cache is capping your batch size)? Out of compute during prefill (long prompts, first-token latency)? Bound by memory bandwidth during decode (each token reads the whole model from HBM)? Or bound by throughput (you have plenty of requests and want more tokens per second per dollar)? Those are four different problems. The rest of this lesson maps tools onto them.
Make the weights smaller: quantization
Decode is memory-bandwidth bound. Generating one token requires streaming every weight from HBM through the compute units. Halve the bytes per weight and you roughly halve the time, and you free memory for a bigger batch or longer context. That is the whole pitch for quantization.
The blunt version is post-training quantization: take trained FP16 weights and round them to INT8 or INT4. Naive rounding wrecks quality because a few outlier weights dominate the error. GPTQ (Frantar et al., 2022) does it smartly. It quantizes weights one column at a time and uses second-order (Hessian) information to compensate the remaining weights for each rounding error, holding 4-bit models close to FP16. AWQ (Lin et al., 2023) takes a different angle: find the small fraction of weight channels that matter most for activations and protect them. On the CPU and laptop end, llama.cpp's GGUF "k-quant" formats (Q4_K_M and friends) do a mixed-precision variant: keep the sensitive layers at higher precision, crush the rest.
What it costs: quality, measured, not assumed. INT8 is nearly free. Well-done 4-bit (GPTQ, AWQ, Q4_K_M) is usually a small perplexity bump most users won't notice. Below 4 bits, degradation gets real and task-dependent. A 2-bit model can look fine on chat and fall apart on code or math. The rule: never ship a quant you haven't evaluated on your task.
Llama-70B storage, rough: FP16 ~140 GB -> needs 2x 80GB GPUs INT8 ~70 GB -> fits one 80GB GPU INT4 ~35 GB -> fits one 48GB GPU, room for KV cache
That table is often the entire reason a deployment is feasible on the hardware you actually have.
Fill the pipe: batching and prefix caching
If quantization is about memory, continuous batching is about not wasting the GPU you already paid for. Old-style static batching groups N requests, runs them together, and waits for the slowest to finish before starting the next batch, so a 500-token answer holds up a 10-token one and the GPU idles. Continuous (a.k.a. in-flight) batching, the idea behind Orca and the default in vLLM and TensorRT-LLM, works at the token level: the moment one sequence emits its stop token, its slot is freed and a queued request drops in. The batch is a living thing that refills every step.
The payoff is throughput, often several times higher at high load. The cost is essentially none for throughput-oriented serving, though a very full batch can nudge up per-request latency, since your request now shares each step with more neighbors.
Prefix caching attacks a different waste: recomputation. Every request with the same system prompt recomputes that prompt's KV cache from scratch. vLLM's PagedAttention stores KV in fixed-size blocks like OS memory pages, which lets it share the blocks for a common prefix across requests. A 2,000-token system prompt sent on every call gets its prefill done once and reused. This is a pure win when prefixes repeat (chat with a fixed persona, few-shot templates, agents that resend a long tool spec) and does nothing when every prompt is unique. Diagnose before you reach for it.
Cut decode latency without touching quality: speculative decoding
Here is the one tool that lowers latency and provably does not change your output distribution. Decode is sequential and bandwidth-bound: one forward pass of a big model per token. Speculative decoding (Leviathan et al., 2022; Chen et al., 2023) breaks the one-token-per-pass rule.
A small, cheap draft model proposes the next k tokens fast. Then the big target model verifies all k in a single forward pass. It can, because scoring tokens it is handed is one parallel pass, unlike generating them one at a time. A modified rejection-sampling step accepts the longest correct prefix and resamples the first wrong token. The math guarantees the accepted tokens are distributed exactly as if the target model had generated them alone.
draft (small) proposes: the cat sat on the target (big) verifies: ok ok ok no -> "a" => 4 tokens confirmed in ONE big-model pass, not 4
When the draft agrees often (routine text, a well-matched draft model), you get roughly 2-3x fewer target passes and a matching latency drop. When it disagrees a lot (hard reasoning, a mismatched draft), you pay for drafts that get thrown away and can end up slower. The cost is a second model in memory and real tuning of k. Variants like Medusa fold the draft into extra prediction heads on the target itself, avoiding the separate model. No quality tradeoff either way. That is the rare part.
When the model won't fit on one GPU: parallelism
Some models don't fit on a single GPU at any precision. Now you split across devices, and the currency becomes inter-GPU communication.
- Tensor parallelism shards each layer's matrices across GPUs; they compute partial results and all-reduce every layer. Low latency, but chatty. It wants a fast interconnect (NVLink), so it lives within a node.
- Pipeline parallelism puts whole layer ranges on different GPUs, passing activations down the line. Cheaper communication, so it can span nodes, but a "bubble" of idle time forms unless you keep many micro-batches in flight.
- Expert parallelism is for Mixture-of-Experts models: spread the experts across GPUs, and route each token to a couple of them. It scales total parameters cheaply but creates uneven, hard-to-predict load as routing shifts.
None of these makes a model faster than fitting it on one GPU would. They are what you do because you can't. The tax is the communication, and you minimize it by keeping tensor-parallel groups on NVLink and pipelining across the slower links between nodes.
The rest of the kit, and how to choose
Prefill/decode disaggregation notices that prefill is compute-heavy and bursty while decode is bandwidth-heavy and steady. Mixing them on one GPU means each interferes with the other's latency. Splitting them onto separate GPU pools (as in DistServe and SGLang) lets you hit tight time-to-first-token and steady inter-token targets independently, at the cost of shipping the KV cache between pools.
CPU/GPU offload (llama.cpp, DeepSpeed-Inference) parks layers or KV in CPU RAM when VRAM runs out. It turns "impossible" into "slow but running," great for a single user on a laptop, a poor trade under load because the PCIe transfer becomes the wall.
Compilation and kernel optimization (fused attention like FlashAttention, CUDA graphs, torch.compile, and what TensorRT-LLM does end to end) squeezes overhead out of the execution itself. Pure upside on latency and throughput, no quality cost; the price is engineering time and flexibility. That is the real spectrum. TensorRT-LLM wrings maximum performance from NVIDIA data-center GPUs at the cost of a rigid build step, while llama.cpp runs a quantized model on nearly anything with a CPU. Neither is "better." They sit at opposite ends of a hardware-and-effort curve, and where you land depends on the bottleneck you actually measured. Measure first, then pick the one tool that moves it.
Sources
- Leviathan, Y., Kalman, M., Matias, Y. "Fast Inference from Transformers via Speculative Decoding." 2022. https://arxiv.org/abs/2211.17192
- Chen, C., et al. "Accelerating Large Language Model Decoding with Speculative Sampling." 2023. https://arxiv.org/abs/2302.01318
- Frantar, E., Ashkboos, S., Hoefler, T., Alistarh, D. "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers." 2022. https://arxiv.org/abs/2210.17323
- Lin, J., et al. "AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration." 2023. https://arxiv.org/abs/2306.00978
- Kwon, W., et al. "Efficient Memory Management for Large Language Model Serving with PagedAttention." SOSP 2023. https://arxiv.org/abs/2309.06180
- Yu, G., et al. "Orca: A Distributed Serving System for Transformer-Based Generative Models." OSDI 2022. https://www.usenix.org/conference/osdi22/presentation/yu
- Zhong, Y., et al. "DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving." OSDI 2024. https://arxiv.org/abs/2401.09670
- NVIDIA. "TensorRT-LLM Documentation." https://nvidia.github.io/TensorRT-LLM/