Infrastructure, Hardware & Production Deployment · 7 min
The Memory Math of Serving a Model
The weights fit in VRAM — so why can't you serve 100 long-context users? Because weights are the small, static part of the bill, and the KV cache, which scales with batch size and sequence length, is what actually blows the budget.
A sysadmin once opened a ticket that read, in full: "24GB card, 14GB model, why does the ninth user OOM the box?" The weights fit with 10GB to spare. The math looked trivial. It wasn't, because the weights are the one part of the bill that never changes, and almost everything expensive scales with the thing he wasn't counting: how many people are talking to the model, and how much they've said.
Here's where the memory actually goes, with numbers you can redo on a napkin.
The part everyone counts: weights
Weight memory is the easy line item because it's static. It's the parameter count times the bytes you store each parameter in.
weight_bytes = num_parameters × bytes_per_parameter fp32 = 4 bytes/param fp16 = 2 bytes/param (bf16 is also 2) int8 = 1 byte/param int4 = 0.5 byte/param
A 7-billion-parameter model in fp16 is 7e9 × 2 = 14 GB. That's it. Quantize to int8 and it's ~7GB; to int4, ~3.5GB. This is the number every model card advertises, and the number that makes people think a 24GB consumer card is roomy. It is roomy — for weights. Weights are furniture. They sit there whether you have one user or a thousand. They don't grow, they don't move, and they are not why your server falls over.
The trap is reading "the weights fit" as "the model fits." Serving a model means holding the weights, plus the state of every conversation in flight, plus the scratch space every forward pass needs. Only the first of those three is fixed.
The part that eats you: the KV cache
To generate token 501 of a reply, the model attends to tokens 1 through 500. Recomputing the keys and values for all 500 prior tokens on every step would be quadratic and absurd, so every serving stack caches them. That cache is the KV cache, and it is per-token, per-layer, per-request. It grows with every token in a conversation and with every conversation you run at once.
Here's the formula. Memorize its shape, not the constant:
kv_bytes = 2 × num_layers × num_kv_heads × head_dim
× seq_len × batch_size × bytes_per_elementThe leading 2 is because you store both keys and values. Everything else is the model's architecture (fixed once you pick the model) multiplied by the two things you control at runtime: seq_len (how long each conversation is) and batch_size (how many you serve at once).
Plug in Llama-2-7B: 32 layers, 32 attention heads, head dimension 128, fp16. Per token:
per_token = 2 × 32 × 32 × 128 × 2 bytes
= 524,288 bytes
≈ 0.5 MB per tokenHalf a megabyte per token sounds harmless until you multiply by a real context. One user holding a 4,096-token conversation:
0.5 MB × 4096 ≈ 2 GB — for one user
Now the ticket answers itself. That 24GB card, minus 14GB of weights, minus a couple GB of overhead we'll get to, leaves roughly 8GB for KV cache. At 2GB per long-context user, that's four users. The ninth user didn't reveal a ninth-user limit — he revealed that the connections already open had finally typed enough to fill the KV budget. Capacity here isn't a headcount; it's a token count, and it tips over the moment the running total of everyone's conversation lengths crosses the line. "Serve 100 long-context users" on this hardware would demand 100 × 2 GB = 200 GB of KV cache alone — fourteen times the size of the model. The weights were never the bill. They were the tip.
Two levers move this number hard, and both are architectural:
Grouped-query attention (GQA). Llama-3-8B has 32 query heads but only 8 key/value heads. The formula keys off num_kv_heads, not query heads, so run the same arithmetic and you get 2 × 32 × 8 × 128 × 2 = 131,072 bytes, about 0.125MB per token — 4× smaller than Llama-2-7B's 0.5MB despite being a bigger model. That single design choice quadruples how many users the same card holds. This is a big reason modern models tolerate long contexts at all — the designers cut the cache, not just the compute.
Cache dtype. Store the KV cache in fp8 or int8 instead of fp16 and you halve it. Many stacks now do this with negligible quality loss, because the cache tolerates lower precision better than the weights do.
The parts the formula forgets
Weights and KV cache are the two big blocks, but a forward pass needs working memory, and two other line items decide whether you actually fit.
Activations and intermediate buffers. Every layer produces intermediate tensors during the forward pass. In inference these are transient — you don't retain the full set the way training does for backprop — but the peak still scales with batch size and sequence length, and it isn't free. The attention scores and the feed-forward expansion (classically 4× the hidden dimension) are the usual peaks. For a large batch of long sequences this can run to a few GB.
Framework overhead. CUDA context, cuBLAS/cuDNN workspaces, NCCL buffers if you're sharding across GPUs, the allocator's own reserved-but-unused pool, and the Python runtime footprint. Budget 1–2GB before you've served a single token. PyTorch's caching allocator in particular holds onto memory it isn't handing back, which is why nvidia-smi often reads higher than your live tensors explain.
A realistic accounting for that 7B model on a 24GB card:
weights (fp16) 14.0 GB framework + CUDA ~1.5 GB peak activations ~1.0 GB ------------------------------------ fixed cost ~16.5 GB KV budget remaining ~7.5 GB → ~3–4 long-context users
Why vLLM changed the answer
The naive way to allocate KV cache is to reserve, per request, enough contiguous memory for the maximum sequence length — because you don't know in advance how long the reply runs. If your max context is 4,096 but the average reply is 300 tokens, you've reserved ten-plus times what you use. The vLLM paper measured 60–80% of KV memory wasted this way on real workloads. That waste is the difference between four users and forty.
PagedAttention, the idea behind vLLM, borrows the operating system's trick for RAM: split the cache into fixed-size blocks and hand them out on demand, like virtual-memory pages. A sequence grows one block at a time; blocks needn't be contiguous; and identical prefixes (the shared system prompt every request carries) can point at the same physical blocks instead of copying. The reported result was 2–4× higher throughput at matched latency, achieved purely by not wasting KV memory. Pair that with continuous batching — swapping a finished request out of the batch and a waiting one in on the very next step, instead of waiting for the slowest request in a fixed batch to finish — and the GPU stays packed with live tokens rather than idling on padding. The lesson under the engineering: KV cache is the binding constraint on how many users you serve, so the winning move is packing it tighter, not buying a bigger card.
What to actually do with this
Before you deploy, compute three numbers, in order:
- Fixed cost — weights plus ~2–3GB overhead. This is your floor.
- Per-token KV cost — run the formula with your model's real
num_kv_headsandnum_layers. Halve it if you cache in fp8. - KV budget —
(VRAM − fixed_cost)divided by per-token cost gives the total tokens you can hold across all live requests at once.
That third number is your true capacity, and you spend it on either many short conversations or few long ones — the product batch × seq_len is what's capped, not the user count. When someone asks "how many users can this box serve," the honest answer is a question back: how long are their conversations? A 512-token support chat and a 100,000-token document analysis run on the same weights and cost roughly 200× different amounts of memory.
Weights tell you whether the model loads. KV math tells you whether the service survives contact with real traffic. Count the second one first.
Sources
- Kwon, W., et al. "Efficient Memory Management for Large Language Model Serving with PagedAttention" (vLLM), 2023. https://arxiv.org/abs/2309.06180
- Vaswani, A., et al. "Attention Is All You Need," 2017. https://arxiv.org/abs/1706.03762
- Ainslie, J., et al. "GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints," 2023. https://arxiv.org/abs/2305.13245
- Shazeer, N. "Fast Transformer Decoding: One Write-Head is All You Need" (multi-query attention), 2019. https://arxiv.org/abs/1911.02150
- Touvron, H., et al. "Llama 2: Open Foundation and Fine-Tuned Chat Models," 2023. https://arxiv.org/abs/2307.09288
- vLLM Documentation, "Conserving Memory" / "Optimization and Tuning." https://docs.vllm.ai/en/latest/
- NVIDIA, "Mastering LLM Techniques: Inference Optimization," 2023. https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/