Infrastructure, Hardware & Production Deployment · 7 min

Deploy and Load-Test: The Module Artifact

The capstone: stand up one model on one runtime, push concurrency until it breaks, and turn TTFT and throughput numbers into a defensible build-vs-buy answer.

A single 8B model on one A10G, one runtime, one config file. That is the whole capstone. You are going to deploy it, point a load generator at it, and turn the graph into a sentence a finance person can act on: "serve it ourselves" or "pay the API." Everything in this module — the memory math, prefill versus decode, the KV cache, the tail latency — collapses into that one decision, and the only honest way to make it is to break your own deployment on purpose and watch where it breaks.

Stand up one model on one runtime

Pick a model and a runtime and stop debating. vLLM is the default for GPU serving because it does continuous batching and PagedAttention out of the box. llama.cpp's server is the default when you are CPU-bound or on a small GPU and want GGUF quantization. Do not run both for the capstone. One model, one runtime, one machine, so the numbers mean something.

# vLLM, one 8B model, one GPU
vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.90 \
  --port 8000

That --gpu-memory-utilization 0.90 is the single most important knob, and it ties straight back to Lesson 1. vLLM grabs 90% of VRAM, subtracts the model weights, and hands the rest to the KV cache. On a 24 GB A10G, an 8B model in FP16 is about 16 GB of weights. That leaves roughly 5.6 GB for KV cache after overhead. Whether that is enough is not a vibe. It is arithmetic you already know how to do.

Before you send a single request, write down the memory budget. KV cache per token is:

bytes_per_token = 2 (K and V)
               * num_layers
               * num_kv_heads * head_dim
               * bytes_per_param (2 for FP16)

Llama-3.1-8B: 32 layers, 8 KV heads (GQA), head_dim 128
= 2 * 32 * 8 * 128 * 2  = 131,072 bytes ≈ 128 KB per token

So 5.6 GB of KV cache holds roughly 45,000 tokens of live context. If your average request is a 1,000-token prompt plus a 500-token generation, that is about 1,500 tokens in flight per request, and you can hold roughly 30 concurrent sequences before the cache is full and vLLM starts queuing or preempting. That number — call it the theoretical concurrency ceiling — is your prediction. The load test is where you find out how wrong it is.

Baseline before you ramp

Never start a load test at load. Start at one. A single request, cold and then warm, gives you the two numbers everything else is measured against: time to first token (TTFT), which is prefill plus queue, and inter-token latency, which is decode.

curl -s -w "\n%{time_starttransfer}s total\n" http://localhost:8000/v1/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"meta-llama/Llama-3.1-8B-Instruct",
       "prompt":"Summarize the CAP theorem in three sentences.",
       "max_tokens":128, "stream":true}'

Run it ten times, throw out the first while weights and CUDA graphs are still warming, and record the median. On an A10G you might see TTFT around 60–90 ms for a short prompt and decode around 30–40 tokens/sec for a single stream. Write these down. When p95 TTFT under load is eight times your baseline, you have found something real; without the baseline you are just staring at a number with no scale.

Use a purpose-built tool for the sweep rather than a bash loop. A curl & fan-out lies to you about concurrency because process spawn time dominates. vLLM ships a serving benchmark (vllm bench serve); the general-purpose option is llmperf from Ray, or wrk/k6 if you only need HTTP-level numbers and are willing to model the token streaming yourself.

# vLLM's own benchmark harness, sweeping request rate
vllm bench serve \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --dataset-name random \
  --random-input-len 1000 --random-output-len 500 \
  --request-rate 4 \
  --num-prompts 200

Run the sweep and find the knee

A concurrency sweep is boring on purpose: hold everything constant, raise one variable, record the same metrics each step. The variable is offered load, in concurrent clients or requests per second. The metrics are TTFT p50/p95, end-to-end p95, aggregate output tokens/sec, and error rate. Step it 1, 2, 4, 8, 16, 32, 64, and keep going until requests start failing or latency goes vertical.

conc   TTFT p50   TTFT p95   e2e p95   tok/s (agg)   errors
  1      75 ms      90 ms     4.1 s        38          0%
  4      95 ms     140 ms     4.4 s       142          0%
  8     130 ms     260 ms     4.9 s       255          0%
 16     210 ms     520 ms     6.2 s       390          0%
 32     480 ms    1,400 ms   11.8 s       470          0%
 48   1,900 ms    5,200 ms   26.0 s       485          2%
 64   4,300 ms   14,000 ms   61.0 s       450         19%

Read that table like a cardiogram. Aggregate throughput climbs, flattens near 32–48 concurrent, then falls at 64 while latency explodes and errors appear. That flattening point is the knee. Below it, adding load buys you throughput almost for free. Above it, you are trading catastrophic latency for no extra work done, because the KV cache is saturated and vLLM is preempting sequences: evicting a half-finished request's cache to admit another, then recomputing it later.

You do not have to guess when that starts. vLLM logs its scheduler state on a timer — running requests, waiting (queued) requests, and GPU KV cache usage as a percentage. The moment you see GPU KV cache usage: 100% and a preemption warning scroll past, you have found the knee's mechanism, not just its symptom. Line the timestamps up against your latency graph and the two will agree. Notice the knee landed near the ~30–48 you predicted from the memory math. When it does not, that gap is the most interesting thing you will learn all week. Maybe your prompts are longer than you assumed, or --max-model-len is capping usable cache, or chunked prefill is stealing decode cycles from in-flight requests.

Define "breaks" before you run so you are not tempted to move the line later. A common SLO: p95 end-to-end under 10 seconds and error rate under 1%. By that rule this deployment's usable ceiling is 32 concurrent, not 48 and definitely not 64. The peak-throughput number and the meets-SLO number are different numbers, and vendors love to quote the first.

Turn latency into a build-vs-buy answer

Now convert the knee into money. At 32 concurrent within SLO, aggregate output is about 470 tokens/sec. Over an hour that is roughly 1.69M output tokens. Price the box: an A10G instance runs about $1.00–$1.20/hr on-demand at the big clouds, cheaper on spot but less reliable. Take $1.10.

self-hosted cost per 1M output tokens (at the SLO knee):
  $1.10/hr  /  (470 tok/s * 3600 s / 1e6)   =  $1.10 / 1.69M  ≈ $0.65 per 1M tokens

Managed API for a comparable open 8B model: ~$0.20 per 1M output tokens

Cost per token is not cost per task, and the capstone asks for the second. A task is one full successful response. At the 32-concurrent knee you were inside SLO with a 0% error rate, so a 500-token answer costs about 500 * $0.65 / 1e6 ≈ $0.00033 — a third of a cent. Push to 48 concurrent and 2% of requests fail; if a failure means a retry, your tokens per successful task climb and the real per-task cost rises faster than the throughput you bought. That is the whole reason you divide spend by successful requests, not offered ones. A deployment that answers 19% of requests with a timeout is not 19% cheaper. It is broken and expensive.

At list prices, buying wins here, and that is the honest, uncomfortable answer the capstone is designed to produce. Self-hosting a single 8B on one GPU only beats the API when utilization is high and sustained: when data cannot leave your network, when you need a fine-tune no API serves, or when you keep the GPU busy around the clock. A box you rent by the hour but use 8 hours a day has its real cost per token tripled, because idle time still bills. Compute your number: hourly_cost / (tokens_per_hour_at_SLO * utilization_fraction). If your utilization is 25%, that $0.65 becomes $2.60 and the API wins by 13x. If you are pinned near full utilization on reserved capacity at $0.40/hr spot, self-hosting drops under the API and building wins.

Both sides of that comparison hide line items outside the token count. The $0.20 API number leaves out rate limits, provider-side cold starts, and the day they reprice. The $0.65 self-host number leaves out the human who keeps the box alive: patching, on-call at 3 a.m., the second GPU you buy for redundancy. Neither omission is a conspiracy. Write them down next to the graph so the decision gets made in daylight, not in a spreadsheet that only counts the easy half.

Make it reproducible or it is a story, not a result. Pin the model revision, the runtime version, the exact flags, the dataset seed, and the hardware SKU in a README next to the results table. Anyone should be able to clone the repo, run one script, and land your knee within noise. That reproducibility is the deliverable. The graph is just its evidence.

Sources