Evaluation methodology
A model page here is a projection of an evaluation database, not a marketing document. This page explains what the numbers on it mean, and the bar a page must clear before we let a search engine index it.
A note while we build: the scores shown on the catalog today are seeded sample data. They exist to show the shape of the page. No harness has been run yet, and the models the platform actually hosts carry no scores at all until one has. Every page below the bar is marked noindex.
A model at tier E may be excellent. The claim is only that nobody has shown it — a different sentence, and the one we are willing to stand behind. The unit is the artifact you would actually run: this fine-tune, at this quantisation. Most of the confusion in this field is a score from one artifact being read as a fact about another.
| Tier | What it takes | What it means for you |
|---|---|---|
A Independent | A third party with no stake in the result published an evaluation of this exact artifact — this fine-tune, this quantisation — with a method reproducible from what they wrote. | Somebody who does not benefit from the answer measured it, and showed their work. |
B Vendor, documented | The maker published numbers AND a technical report or paper describing how they were produced. Not independently reproduced. | You can see the method and argue with it. You are still taking the measurement from the party it flatters. |
C Vendor, undocumented | Numbers on a model card or a launch post, with no method, no report, and no reproduction. | A figure with nothing behind it. You cannot check it, and neither can anyone else. |
D Proxy only | The only numbers in the chain belong to something else — the base model, a sibling checkpoint, an earlier version, or the full-precision weights when the artifact you would run is quantised. | Real evidence, about a different artifact. The most common mistake in this field is treating it as evidence about this one. |
E None | No capability numbers exist anywhere in the chain. Marketing claims, benchmark logos, and community enthusiasm are not numbers. | Nobody has measured it. That is not an accusation; it is the state of the record. |
A compliance rate is not a capability score. A refusal or compliance rate has been published. It measures willingness to answer, not whether the answer is right — a witness who answers every question is not thereby a reliable witness. So it is flagged rather than lettered: a model whose only published number is a refusal rate is a tier E on capability, carrying a note that its willingness has been measured.
When the record is empty we say so with a date and the search behind it — where we looked, what we searched for, and when. A negative claim without a method is an assertion; with one it is a finding somebody can correct. That applies to us too: where a page says a number is unmeasured, it includes unmeasured by us.
Each row is one metric from one evaluation suite, run against one model version. Scores are re-computed, never edited: the results table is append-only, so a metric’s current value is simply its newest run. Older runs stay readable as history. A version does not inherit another version’s numbers.
A number on its own is not evidence. Each qualifying row must carry, in the same table, all of:
- A named baseline — the strongest general-purpose model we could run on the same suite, the same setup, the same day. A score is only interpretable against what a competent generalist scored on the identical task.
- A sample size (n) — a score over nine items is not a measurement. n is shown so you can judge it.
- A harness version — the exact evaluation code that produced the run, so a number is reproducible and comparable across models.
- A run date — a score is a day-grained fact, not a permanent claim.
- A negative control — at least one published row where the model is expected to do badly — and does. The control is what tells you the other rows are worth anything; it is rendered inline and flagged, never hidden in a footnote.
Indexability is derived, never set by hand. A model page is served for anyone with its URL, but it is only listed in the catalog and allowed into a search index once it clears the bar:
- at least 3 complete results,
- across at least 2 distinct suites (three results in one suite is one experiment reported three ways),
- including at least one non-control result, and
- a published negative control.
There is no override that forces a thin page into the index. If we want a page indexed, we run the eval. An operator can force a page out of the index, never in.
We do not invent scores, and we do not name a competitor we did not actually run — a made-up baseline is a defamatory-grade claim, not a measurement. A model with no run has no scores and says so (“not scored yet”), rather than borrowing a believable-looking number. That empty state is the honest one.
Questions about the methodology? Reach AdversariaLLM at the contact page.