All field notes

Guide · Foundations · D. Rose · 5 August 2026 · Updated 13 August 2026 · 4 min

Base, instruct, open, closed: reading a model’s label

What the words on a model card mean, and why parameter count is the least interesting number on it.

Pick any model and you meet a wall of labels: base or instruct, 7B or 70B, open or closed weights, dense or mixture-of-experts, maybe a "-chat" or "-it" suffix. Each tells you something real about how the thing will behave and how you are allowed to use it. Reading them is a skill worth having before you compare a single benchmark.

Base versus instruct is the distinction that matters most and confuses newcomers the most. A base model is the raw result of training on a mountain of text: it does nothing but continue whatever you give it. Hand it "The capital of France is" and it completes "Paris." Hand it a question and it might answer, or might continue with three more questions, because a list of questions is also a plausible continuation. A base model has no notion that it is supposed to be helpful. An instruct (or "chat", "-it") model is a base model trained further to follow instructions and hold a conversation — to treat your text as a request to satisfy rather than a passage to extend. Almost every model you interact with as an assistant is an instruct model. The base model is the foundation underneath it, still useful, but a different tool with a different interface. Confusing the two is why someone occasionally reports that a model "just rambled instead of answering" — they were talking to a base model.

Open versus closed is about the weights, not the price. A closed model — the ones behind the big commercial APIs — you reach only over the network; the weights never leave the provider, you cannot inspect or run them, and your data goes to their servers under their terms. An open-weights model has its numbers published: you can download them, run them on your own hardware, fine-tune them, and inspect what you deploy. "Open" does a lot of quiet work here — many open-weights models carry licenses that restrict commercial use or redistribution, and few publish their training data — so "open weights" is not the same as "open source," and reading the actual license matters, especially in security contexts where where the data goes is part of the threat model.

Now the number everyone fixates on: parameters. 7B means roughly seven billion learned numbers; 70B, ten times that. More parameters mean more capacity to absorb patterns, and, all else equal, a bigger model tends to be more capable. But all else is never equal, and parameter count is a poor single predictor of quality, for a few concrete reasons.

The first is data. A model's ability comes as much from how much good text it was trained on as from its size. The well-known result here — associated with the "Chinchilla" work — is that many early large models were badly undertrained: for a fixed compute budget you are better off with a smaller model fed more data than a giant model fed too little. A well-trained 8B model routinely beats a poorly-trained model several times its size. Size is potential; training data is how much of it got realized.

The second is what the parameters are for. A mixture-of-experts (MoE) model may advertise a huge total parameter count but only activate a small fraction of it for any given token — it routes each token to a couple of specialized "experts" rather than running the whole network. So a headline "total parameters" number and the "active parameters" that actually run per token can differ by an order of magnitude, and only the second tracks the cost and latency of running it. When you see a very large MoE model, ask which number is being quoted.

The third is fit. Capability is not one axis. A small model fine-tuned for detection engineering or for reading disassembly can outperform a far larger general model on exactly that task while losing to it everywhere else. "Which model is best" is not a question with an answer; "best for this task, at this latency, at this cost, running where my data is allowed to go" is. A small specialized model you run locally, a large general model behind an API, and a plain deterministic script are all valid answers, and the interesting engineering is choosing deliberately rather than reaching for the biggest number.

You will also meet packaging and precision labels — GGUF, GPTQ, AWQ, "4-bit" — and it is worth not conflating them. GGUF is a file format for storing and running a model, a container. GPTQ and AWQ are quantization methods: ways of squeezing the weights down to fewer bits so the model fits in less memory and runs faster, at some cost to quality. One is about how the model is packaged; the other about how much precision was traded away to shrink it. Both sit on the card next to the parameter count, and both tell you more about what it will actually cost you to run than the parameter count does.

More in Foundations