Guide · Foundations · D. Rose · 5 September 2026 · 4 min
Why a base model's scores are a ceiling
Fine-tuning is not addition — the weights holding the old capability are the ones being changed. Catastrophic forgetting is why every model page treats the base benchmark as a best case.
Nearly every model in this catalog is a fine-tune, and nearly every model page carries a version of the same sentence: the published numbers belong to the base model, so treat them as a ceiling.
A ceiling, not an estimate. That word is doing real work — it says a fine-tune can be worse than the model it started from, at things the fine-tune was never about. This is the page that explains why that is the honest default rather than a pessimistic one.
Fine-tuning is not addition
The intuitive model of a security fine-tune is arithmetic: general model, plus security knowledge, equals security model. Nothing is subtracted, so nothing is lost.
That is not what happens. Fine-tuning is more training on the same weights. Gradients from the new data move parameters toward the new objective, and those parameters are the same ones holding everything the model already knew. There is no separate compartment for the old capability to sit safely in — the weights that encode how to write a coherent paragraph are the weights being adjusted to make it better at CVE triage.
Where the new task and the old capability want different values, the new task wins, because it is the only one with a loss function present in the room. Anything the fine-tuning data does not exercise has nothing defending it.
The field's name for the result is catastrophic forgetting, and it long predates language models.
It is measured, and it gets worse with size
This is not a theoretical worry. A 2023 study evaluated forgetting during continual instruction tuning across domain knowledge, reasoning and reading comprehension, and found that
catastrophic forgetting is generally observed in LLMs ranging from 1b to 7b parameters
with a second finding that matters more for this catalog than the first: as the models got larger within that range, the forgetting got worse, not better — the authors attribute it to the larger model having had more to lose.
A 2024 paper at EMNLP went after the cause and tied the extent of forgetting to the flatness of the loss landscape around the fine-tuned solution, then used that to mitigate it. Their framing of the state of play is worth quoting, because it is a paper about a well-known effect saying the mechanism was still open:
It compromises the effectiveness of large language models (LLMs) during fine-tuning, yet the underlying causes have not been thoroughly investigated.
Be precise about what those cover, though. Both study particular setups — continual instruction tuning, specific scales, specific datasets. Neither one licenses a number for a 33B security fine-tune by an anonymous account. What they establish is direction and plausibility, not magnitude.
Why that makes the base score a ceiling
Put the two halves together and the catalog's phrasing follows.
A fine-tune might have gained the thing it claims — nobody measured it. And it might have lost general capability along the way — nobody measured that either. The base model's benchmark is a number for an artifact that has since been modified in a direction known to cost something, by an unknown amount, on axes nobody re-tested.
That is why the number is a ceiling and not a description. It is the best case, and it is a best case that requires the fine-tuning to have been free, which the literature says it generally is not.
None of this says fine-tunes are bad. A security fine-tune can be enormously more useful than its base on the work you care about, and the trade is often worth making. The point is narrower: the trade is real, and the numbers on the page are from before it was made.
Where it compounds
Most models here are not one edit away from their base. A typical community build is a fine-tune of an abliterated model, then quantised to four bits. Three transformations, each with an unmeasured cost, and a benchmark table describing the state before any of them.
That is tier D in one sentence: real evidence, about a different artifact.
What would settle it
Nothing here is unmeasurable. The fine-tune could be run against a general benchmark it was not trained for and compared with its base — the same evaluation, both artifacts, published together. That is a day of compute and it is the difference between "we improved it" and "we improved it and here is what it cost".
Almost no community model card does this. When one does, it earns a better grade on this site, and it should.
Sources. Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou and Yue Zhang, "An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning", 2023. Hongyu Li, Liang Ding, Meng Fang and Dacheng Tao, "Revisiting Catastrophic Forgetting in Large Language Model Tuning", Findings of the ACL: EMNLP 2024. Both quotations are from the papers' abstracts. Retrieved 5 September 2026.