Infrastructure, Hardware & Production Deployment · 7 min
Managed API Versus Self-Hosting: An Engineering Decision
Neither path is intrinsically cheaper or safer. The honest comparison turns on four hinges — utilization, control, contracts, and your operating model — and utilization is where the money crosses over.
A company I know ran the same 7-billion-parameter model two ways in the same quarter. One team paid a hosted API per token. The other rented a GPU by the hour and served it themselves. At the end of the quarter the API team had spent less money and shipped faster. Six months later, after traffic grew 20x and stabilized, the self-hosted team was spending a third of what the API path would have cost at that volume. Nobody was wrong. The right answer moved because the workload moved.
That is the whole argument in miniature. "Managed API versus self-hosting" gets argued as if one side is intrinsically cheaper, or one side is intrinsically more secure. Both claims are marketing. The honest version is an engineering tradeoff with a small number of real hinges: utilization, control, contracts, and how much operations work you can actually absorb.
The two cost models are shaped differently
A managed API charges you per unit of work: dollars per million input tokens, dollars per million output tokens, sometimes a surcharge for cached context or long outputs. Your bill is a function of what you use. Send nothing over a weekend and you pay nothing.
Self-hosting charges you for capacity, whether or not you use it. An H100 on a cloud runs roughly $2 to $4 per hour on demand at 2025 rates, cheaper on committed or spot pricing. That meter runs at 3 a.m. when your traffic is zero. Your bill is a function of how long the machine is on, not how hard it works.
That difference is the entire break-even argument, and you can write it as one inequality. Let your GPU cost C dollars per hour and deliver a sustained throughput of T tokens per second when busy. Its cost per token at full utilization is:
cost_per_token = C / (T * 3600)
Put numbers in. A single H100 serving a 7B model with a decent inference server can push well past 2,000 output tokens/second aggregated across concurrent requests. At $3/hour:
3 / (2000 * 3600) = $0.00000042 per token ≈ $0.42 per million tokens
Hold that against a menu price. Hosted models commonly list somewhere between $0.15 and a few dollars per million input tokens, with output tokens costing several times more; frontier models sit at the top of that range and change often, so price from the vendor's live page, not from memory. Against those numbers a saturated H100 looks crushingly cheap. But that number assumes the GPU is saturated. It almost never is.
Utilization is the hinge
Rewrite the formula with a utilization factor u, the fraction of wall-clock time the GPU is doing useful work:
effective_cost_per_token = C / (T * 3600 * u)
At u = 1.0 you get $0.42/M. At u = 0.1 — a real number for a service with spiky daytime traffic and no autoscaling — you get $4.20/M, and now the managed API is competitive or cheaper, and it came with zero operational burden.
This is why the API team in the opening story won early. A new product has bursty, unpredictable, low-average traffic. You are paying for a full GPU to sit idle between requests. The per-token math on paper says self-hosting is cheap; the per-token math you actually pay says otherwise, because the denominator has a u in it and your u is 0.08.
Utilization climbs when traffic is high, steady, and batchable. A backfill job re-summarizing ten million documents overnight can drive u toward 1.0, because you can keep the batch queue full. That is the workload where self-hosting's capacity model wins outright. So the first question is not "which is cheaper" but "what does my traffic shape look like, and can I keep the hardware busy?"
A rough decision sketch:
low/spiky/unpredictable volume -> managed API (you can't fill the GPU) high/steady/batchable volume -> self-host (you can, and per-token plummets) somewhere in between -> measure u before you commit capital
What you are actually buying on each side
Cost is only one axis. The others decide as many real cases.
Infrastructure burden versus runtime control. With an API you inherit none of the pager. No CUDA driver mismatches, no OOM at 2,000 tokens of context, no capacity planning, no model-loading cold starts. You also get no control over the runtime. You cannot pin a kernel, change the batching policy, swap the attention implementation, or run a quantization the provider does not offer. Self-hosting hands you all of that control and all of that pager.
Frontier access versus customization. The strongest general models are available only through managed APIs — you cannot rent the weights for a top proprietary model at any price. If your product needs that capability today, the API is not a preference, it is the only door. Self-hosting's counter-argument is that you can run exactly the model and runtime you want: a fine-tuned checkpoint, an open-weights model with a custom LoRA, a specific quantization, a modified serving stack. You trade the frontier for the ability to shape the thing.
Provider-defined boundary versus custom isolation. This is where the "self-hosting is more secure" dogma dies. A serious managed provider gives you a real, contractual trust boundary: SOC 2 reports, data-retention and training-use terms, tenant isolation you can point an auditor at. Most teams cannot build isolation that good on their own. What self-hosting gives you is a different boundary — one where the data never leaves a network you control, which is not "more secure" in the abstract but is decisive when a regulation or contract says the bytes may not cross a third party's wire. Read the provider's data-use terms; that document, not a vibe, is your security posture on the managed path.
Less tuning versus more optimization responsibility. The API abstracts prefill, decode, KV-cache management, and batching away from you. That is less to tune and less you can tune. Self-hosting makes throughput your job: you own the batch size, the KV-cache memory budget, the max context, and the GPU-fitting math from earlier in this module. That is leverage if you have the skill and staff, and a liability if you don't.
How to actually decide
Do not decide from a blog post's opinion, including this one. Decide from your own numbers.
- Estimate steady-state volume in tokens/day, and its shape over the day. This is the single most predictive input.
- Price the managed path:
tokens/day * blended_price_per_token. Use a blended rate that reflects your real input/output mix; output tokens usually cost several times more than input. - Price the self-host path honestly. Not just GPU-hours. Add engineer time to build and run the serving stack, on-call, and — critically — your realistic
u. A common mistake is pricing atu = 1.0. - Find the crossover. There is a tokens/day figure above which self-hosting wins and below which it loses. Know where you sit relative to it, and how fast you're moving toward it.
def monthly_cost(tokens_per_day, api_price_per_tok, gpu_hourly,
tok_per_sec_busy, utilization, gpus=1):
api = tokens_per_day * 30 * api_price_per_tok
gpu_tokens_per_month = tok_per_sec_busy * utilization * 3600 * 24 * 30 * gpus
if gpu_tokens_per_month < tokens_per_day * 30:
return api, float("inf") # can't even serve the load
self_host = gpu_hourly * 24 * 30 * gpus
return api, self_hostRun that with your numbers, sweep utilization from 0.05 to 0.9, and watch the crossover move. That sweep, not a doctrine, is the answer.
Anchor the sweep with one number. An always-on $3/hour H100 costs about $2,160 a month. At a blended $0.50 per million tokens, a managed API charges you that same $2,160 only once you are pushing on the order of 140 million tokens a day. Below that the API is simply cheaper, full stop. And serving 140 million tokens a day off a single H100 already demands utilization north of 0.8 — which is why, in practice, the crossover point and the saturation point tend to arrive together.
The mature pattern is rarely all-or-nothing. Plenty of teams start on a managed API to ship, instrument their real traffic, and move only the high-volume, steady, non-frontier workloads in-house once u justifies the capital and the headcount — while leaving spiky or frontier-dependent traffic on the API. Neither path is a creed. They are tools, and the workload picks the tool.
Sources
- Kwon, Woosuk, et al. "Efficient Memory Management for Large Language Model Serving with PagedAttention." SOSP, 2023. https://arxiv.org/abs/2309.06180
- vLLM Project. "vLLM Documentation." 2024. https://docs.vllm.ai/
- NVIDIA. "H100 Tensor Core GPU Datasheet." 2023. https://resources.nvidia.com/en-us-tensor-core/nvidia-tensor-core-gpu-datasheet
- Amazon Web Services. "AWS Well-Architected Framework — Cost Optimization Pillar." 2023. https://docs.aws.amazon.com/wellarchitected/latest/cost-optimization-pillar/welcome.html
- OpenAI. "API Pricing." 2024. https://openai.com/api/pricing/
- Anthropic. "Pricing." 2024. https://www.anthropic.com/pricing
- AICPA & CIMA. "System and Organization Controls (SOC) Suite of Services." 2022. https://www.aicpa-cima.com/topic/audit-assurance/audit-and-assurance-greater-than-soc-2