Infrastructure, Hardware & Production Deployment · 7 min
Deploy and Load-Test: The Module Artifact
The capstone: stand up one model on one runtime, push concurrency until it breaks, and turn TTFT and throughput numbers into a defensible build-vs-buy answer.
A single 8B model on one A10G, one runtime, one config file. That is the whole capstone. You are going to deploy it, point a load generator at it, and turn the graph into a sentence a finance person can act on: "serve it ourselves" or "pay the API." Everything in this module — the memory math, prefill versus decode, the KV cache, the tail latency — collapses into that one decision, and the only honest way to make it is to break your own deployment on purpose and watch where it breaks.
Stand up one model on one runtime
Pick a model and a runtime and stop debating. vLLM is the default for GPU serving because it does continuous batching and PagedAttention out of the box. llama.cpp's server is the default when you are CPU-bound or on a small GPU and want GGUF quantization. Do not run both for the capstone. One model, one runtime, one machine, so the numbers mean something.
# vLLM, one 8B model, one GPU vllm serve meta-llama/Llama-3.1-8B-Instruct \ --max-model-len 8192 \ --gpu-memory-utilization 0.90 \ --port 8000
That --gpu-memory-utilization 0.90 is the single most important knob, and it ties straight back to Lesson 1. vLLM grabs 90% of VRAM, subtracts the model weights, and hands the rest to the KV cache. On a 24 GB A10G, an 8B model in FP16 is about 16 GB of weights. That leaves roughly 5.6 GB for KV cache after overhead. Whether that is enough is not a vibe. It is arithmetic you already know how to do.
Before you send a single request, write down the memory budget. KV cache per token is:
bytes_per_token = 2 (K and V)
* num_layers
* num_kv_heads * head_dim
* bytes_per_param (2 for FP16)
Llama-3.1-8B: 32 layers, 8 KV heads (GQA), head_dim 128
= 2 * 32 * 8 * 128 * 2 = 131,072 bytes ≈ 128 KB per tokenSo 5.6 GB of KV cache holds roughly 45,000 tokens of live context. If your average request is a 1,000-token prompt plus a 500-token generation, that is about 1,500 tokens in flight per request, and you can hold roughly 30 concurrent sequences before the cache is full and vLLM starts queuing or preempting. That number — call it the theoretical concurrency ceiling — is your prediction. The load test is where you find out how wrong it is.
Baseline before you ramp
Never start a load test at load. Start at one. A single request, cold and then warm, gives you the two numbers everything else is measured against: time to first token (TTFT), which is prefill plus queue, and inter-token latency, which is decode.
curl -s -w "\n%{time_starttransfer}s total\n" http://localhost:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{"model":"meta-llama/Llama-3.1-8B-Instruct",
"prompt":"Summarize the CAP theorem in three sentences.",
"max_tokens":128, "stream":true}'Run it ten times, throw out the first while weights and CUDA graphs are still warming, and record the median. On an A10G you might see TTFT around 60–90 ms for a short prompt and decode around 30–40 tokens/sec for a single stream. Write these down. When p95 TTFT under load is eight times your baseline, you have found something real; without the baseline you are just staring at a number with no scale.
Use a purpose-built tool for the sweep rather than a bash loop. A curl & fan-out lies to you about concurrency because process spawn time dominates. vLLM ships a serving benchmark (vllm bench serve); the general-purpose option is llmperf from Ray, or wrk/k6 if you only need HTTP-level numbers and are willing to model the token streaming yourself.
# vLLM's own benchmark harness, sweeping request rate vllm bench serve \ --model meta-llama/Llama-3.1-8B-Instruct \ --dataset-name random \ --random-input-len 1000 --random-output-len 500 \ --request-rate 4 \ --num-prompts 200
Run the sweep and find the knee
A concurrency sweep is boring on purpose: hold everything constant, raise one variable, record the same metrics each step. The variable is offered load, in concurrent clients or requests per second. The metrics are TTFT p50/p95, end-to-end p95, aggregate output tokens/sec, and error rate. Step it 1, 2, 4, 8, 16, 32, 64, and keep going until requests start failing or latency goes vertical.
conc TTFT p50 TTFT p95 e2e p95 tok/s (agg) errors 1 75 ms 90 ms 4.1 s 38 0% 4 95 ms 140 ms 4.4 s 142 0% 8 130 ms 260 ms 4.9 s 255 0% 16 210 ms 520 ms 6.2 s 390 0% 32 480 ms 1,400 ms 11.8 s 470 0% 48 1,900 ms 5,200 ms 26.0 s 485 2% 64 4,300 ms 14,000 ms 61.0 s 450 19%
Read that table like a cardiogram. Aggregate throughput climbs, flattens near 32–48 concurrent, then falls at 64 while latency explodes and errors appear. That flattening point is the knee. Below it, adding load buys you throughput almost for free. Above it, you are trading catastrophic latency for no extra work done, because the KV cache is saturated and vLLM is preempting sequences: evicting a half-finished request's cache to admit another, then recomputing it later.
You do not have to guess when that starts. vLLM logs its scheduler state on a timer — running requests, waiting (queued) requests, and GPU KV cache usage as a percentage. The moment you see GPU KV cache usage: 100% and a preemption warning scroll past, you have found the knee's mechanism, not just its symptom. Line the timestamps up against your latency graph and the two will agree. Notice the knee landed near the ~30–48 you predicted from the memory math. When it does not, that gap is the most interesting thing you will learn all week. Maybe your prompts are longer than you assumed, or --max-model-len is capping usable cache, or chunked prefill is stealing decode cycles from in-flight requests.
Define "breaks" before you run so you are not tempted to move the line later. A common SLO: p95 end-to-end under 10 seconds and error rate under 1%. By that rule this deployment's usable ceiling is 32 concurrent, not 48 and definitely not 64. The peak-throughput number and the meets-SLO number are different numbers, and vendors love to quote the first.
Turn latency into a build-vs-buy answer
Now convert the knee into money. At 32 concurrent within SLO, aggregate output is about 470 tokens/sec. Over an hour that is roughly 1.69M output tokens. Price the box: an A10G instance runs about $1.00–$1.20/hr on-demand at the big clouds, cheaper on spot but less reliable. Take $1.10.
self-hosted cost per 1M output tokens (at the SLO knee): $1.10/hr / (470 tok/s * 3600 s / 1e6) = $1.10 / 1.69M ≈ $0.65 per 1M tokens Managed API for a comparable open 8B model: ~$0.20 per 1M output tokens
Cost per token is not cost per task, and the capstone asks for the second. A task is one full successful response. At the 32-concurrent knee you were inside SLO with a 0% error rate, so a 500-token answer costs about 500 * $0.65 / 1e6 ≈ $0.00033 — a third of a cent. Push to 48 concurrent and 2% of requests fail; if a failure means a retry, your tokens per successful task climb and the real per-task cost rises faster than the throughput you bought. That is the whole reason you divide spend by successful requests, not offered ones. A deployment that answers 19% of requests with a timeout is not 19% cheaper. It is broken and expensive.
At list prices, buying wins here, and that is the honest, uncomfortable answer the capstone is designed to produce. Self-hosting a single 8B on one GPU only beats the API when utilization is high and sustained: when data cannot leave your network, when you need a fine-tune no API serves, or when you keep the GPU busy around the clock. A box you rent by the hour but use 8 hours a day has its real cost per token tripled, because idle time still bills. Compute your number: hourly_cost / (tokens_per_hour_at_SLO * utilization_fraction). If your utilization is 25%, that $0.65 becomes $2.60 and the API wins by 13x. If you are pinned near full utilization on reserved capacity at $0.40/hr spot, self-hosting drops under the API and building wins.
Both sides of that comparison hide line items outside the token count. The $0.20 API number leaves out rate limits, provider-side cold starts, and the day they reprice. The $0.65 self-host number leaves out the human who keeps the box alive: patching, on-call at 3 a.m., the second GPU you buy for redundancy. Neither omission is a conspiracy. Write them down next to the graph so the decision gets made in daylight, not in a spreadsheet that only counts the easy half.
Make it reproducible or it is a story, not a result. Pin the model revision, the runtime version, the exact flags, the dataset seed, and the hardware SKU in a README next to the results table. Anyone should be able to clone the repo, run one script, and land your knee within noise. That reproducibility is the deliverable. The graph is just its evidence.
Sources
- Kwon, Woosuk et al. "Efficient Memory Management for Large Language Model Serving with PagedAttention." SOSP 2023. https://arxiv.org/abs/2309.06180
- Ainslie, Joshua et al. "GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints." 2023. https://arxiv.org/abs/2305.13245
- vLLM Project. "vLLM Documentation." 2023–2025. https://docs.vllm.ai/en/latest/
- vLLM Project. "vLLM benchmarks (benchmark_serving)." GitHub, 2023–2025. https://github.com/vllm-project/vllm/tree/main/benchmarks
- ggml-org. "llama.cpp — LLM inference in C/C++." 2023–2025. https://github.com/ggml-org/llama.cpp
- Ray Project. "LLMPerf: A tool for evaluating the performance of LLM APIs." 2024. https://github.com/ray-project/llmperf
- Grafana Labs. "k6 Documentation." 2024. https://grafana.com/docs/k6/latest/