Model Families & Architectures · 6 min

Open-Weight vs Closed Families: What It Means in Practice

Open-weight and closed API models differ in data control, licensing, fine-tuning access, and cost shape, and that sourcing choice is separate from runtime routing.

Two questions get collapsed into one all the time, and it drives bad architecture. The first is where does this model come from: do you hold the weights, or does someone else? The second is which model answers this particular request at runtime. They feel like the same question ("should we use Llama or GPT?"), but they are orthogonal. You can self-host an open-weight model and still route your hardest requests to a closed API from the same service. Get the sourcing axis right first. Routing is a separate lesson.

The mental model: what you actually possess

A closed family (GPT, Claude, Gemini) is a capability you rent through a socket. You send tokens over HTTPS, you get tokens back, and you never touch the parameters. An open-weight family (Llama, Mistral/Mixtral, Qwen, and friends) is a several-gigabyte file of floating-point numbers you can download, load onto a GPU you control, and run offline forever.

That one difference, whether the weights sit on your disk, cascades into everything practical: where your data goes, what you are legally allowed to do, whether you can fine-tune, and the shape of your bill.

CLOSED (rented)              OPEN-WEIGHT (possessed)
  your prompt                   your prompt
      │                             │
      ▼                             ▼
 [provider's GPUs] ← data      [your GPU / your VPC]
      │             leaves          │  data never leaves
      ▼             your boundary   ▼
   tokens back                  tokens back
  pay per token                 pay per GPU-hour
  no weights                    weights on your disk

Data control is the sharpest edge

With a closed API, your prompts and outputs traverse someone else's infrastructure. Providers now offer zero-retention and no-training-on-your-data enterprise terms, and those are real and contractual. But they are still terms, not physics. When your threat model or your regulator will not accept "trust the vendor's retention policy," self-hosting an open-weight model is the only answer that keeps every token inside a boundary you own. An air-gapped deployment is possible only with open weights. Through a socket it is impossible by definition.

Defender: "Data control" is not binary. A self-hosted model still logs, caches KV state, and can leak through your own telemetry. Owning the weights moves the trust boundary to you, which means the audit burden moves to you too. Do not confuse "on-prem" with "secure by default."

"Open weight" is not "open source"

This trips up procurement constantly. Open weights means you got the final trained parameters. It does not mean you got the training code or the dataset, and without those you cannot reproduce, fully audit, or explain the model's behavior from first principles. The OSI's Open Source AI Definition (OSAID 1.0, October 2024) sets a genuinely higher bar: freedom to use, study, modify, and share, backed by enough data information that a skilled person could rebuild an equivalent model. Most "open" LLMs, Llama included, clear the "you can run it" bar and fail the OSAID bar.

Researcher: For reproducible science you need weights plus data provenance plus checkpoints. Open weights alone let you probe behavior but not attribute it: you cannot tell whether a capability came from data, architecture, or RLHF. If your paper's claim rests on that distinction, an open-weight-only model is a black box wearing an open-source hat.

Licenses: read them, they bite

The failure mode is assuming "open weight" means "do whatever." It does not.

Llama ships under Meta's Community License, which is source-available, not open source. Two things bite in practice. First, a scale cliff: if your product crosses 700 million monthly active users, you have to go negotiate a separate license with Meta. Second, attribution is mandatory: you must display "Built with Llama," and any model you train on Llama or on its outputs must be named with a "Llama" prefix. Watch what changed here, because stale advice is everywhere. Llama 2 and Llama 3 flatly prohibited using Llama's outputs to improve any competing model. Llama 3.1 dropped that prohibition. You may now train on synthetic Llama data; you just have to name the result "Llama-something" and keep the attribution. Cite the version, not the folklore.

Mistral is genuinely tiered. Mistral 7B and Mixtral 8x7B / 8x22B are Apache 2.0: real permissive open source, no attribution gymnastics, no user-count cliff. But flagship and specialty models ship under the Mistral Research License (non-commercial) or the Non-Production License (Codestral, for one). Same vendor, three completely different sets of rights. Never reason about "the Mistral license" as if it were one thing.

Builder: Pin the license to the exact checkpoint in your lockfile, not to the model family. "We use Mixtral" tells your lawyer nothing; "mixtral-8x7b-v0.1, Apache-2.0" tells them everything. A model swap can silently change your legal footing.

Fine-tuning access is asymmetric

Want to bake domain behavior into the weights? Open-weight models give you the full spectrum: LoRA adapters, full fine-tunes, quantization, distillation, whatever your GPUs can handle. Closed families deliberately narrow this. OpenAI exposes a hosted fine-tuning API. Anthropic offers no general public fine-tuning for Claude through its own API; the nearest thing is Claude 3 Haiku fine-tuning inside Amazon Bedrock (GA since November 2024), plus enterprise custom-training deals. Google offers tuning through Vertex AI. So "we need to fine-tune on proprietary data" is often the deciding constraint that pushes a workload onto open weights, or onto the one closed vendor that happens to expose it.

Cost is a different shape, not just a different number

Closed APIs are pay-per-token with zero idle cost. GPT-4o runs roughly $2.50 per million input tokens and $10 per million output (down from $5/$15 at its 2024 launch); GPT-4o-mini is about $0.15/$0.60. You pay for exactly what you use and nothing while idle.

Self-hosting is pay-per-GPU-hour, and idle time is pure waste. Run the arithmetic before you believe "open is cheaper." A mid-tier inference GPU rents for roughly $1/hour, about $720/month if you keep it up. At GPT-4o-mini's $0.60 per million output tokens, that same $720 buys about 1.2 billion output tokens from the API. So a self-hosted 8B model wins on raw price only if you are generating north of a billion tokens a month and keeping the GPU busy enough to amortize it. A box sitting at 5% utilization is the most expensive model you own.

The honest decision rule:

low / bursty volume ........ closed API wins (no idle cost)
high, steady volume ........ self-host can win (amortized GPU)
data can't leave boundary .. self-host, price be damned
need to fine-tune weights .. open (or the one vendor that allows it)
need frontier reasoning .... closed, today

Back to the two axes

Sourcing (own vs rent) and routing (which model per request) are independent. A mature system usually does both. It self-hosts an open-weight model for the 90% of cheap, high-volume, privacy-sensitive traffic, and routes the 10% of genuinely hard requests to a closed frontier API, with a policy layer deciding per request. Choosing which model handles a given call at inference time is exactly the runtime routing problem, and it is the next thing to get right once your sourcing mix is settled.

Sources

Open-Weight vs Closed Families: What It Means in Practice — All About LLMs, from AI to Z · AdversariaLLM