Guide · Foundations · D. Rose · 5 September 2026 · 4 min
What quantization actually costs
4-bit takes about three-tenths of a point off the average and two-thirds off the file. The average is also the part that hides what changed — and none of the numbers describe the model you are running.
Every model this catalog serves is quantised. The file you actually run is not the file the maker trained, and on nine of these research posts you will find a sentence noting that the published numbers describe the full-precision weights instead. That sentence deserves a page of its own, because "quantised" gets read either as "basically the same" or as "ruined", and the measured answer is neither.
What the compression actually is
A trained model is a very large pile of numbers. Released weights are usually 16-bit floats — two bytes per parameter, so an 8B model is about 16 GB before anything else is loaded.
Quantisation stores those numbers with fewer bits. Not by rounding each one independently, which would be crude, but block-wise: weights are grouped, each group gets a shared scale factor, and the individual values become small integers relative to that scale. The llama.cpp "K-quant" formats go further and spend different precision on different tensors, which is why Q4_K_M is not uniformly four bits and why the letters on the end are not decoration.
Nothing is learned or added. It is lossy compression of numbers that already existed, chosen so the loss lands where the model is least sensitive.
What it measurably costs
Somebody ran the whole grid. Llama-3.1-8B-Instruct, converted to GGUF, thirteen llama.cpp quantisations against the FP16 original, on five benchmarks plus perplexity. Averaging across the five:
| Build | Size cut | Avg score | Perplexity |
|---|---|---|---|
| FP16 (baseline) | — | 69.47 | 7.32 |
| Q8_0 | 46.87% | 69.41 | 7.33 |
| Q4_K_M | 69.41% | 69.15 | 7.56 |
| Q3_K_S | 77.23% | 65.49 | 8.96 |
Read the Q4_K_M row carefully, because it is the one most people actually download. Two-thirds of the file is gone and the average moves by about three-tenths of a point. On that evidence the fear is misplaced: 4-bit is not where models break.
Three bits is a different story — four points off the average, perplexity up by more than a point and a half, and the damage concentrated in the multi-step reasoning task rather than spread evenly. If you are choosing on footprint alone, that is the boundary worth knowing about.
What the average is hiding
Now the part that matters more than the table.
In the same study, the grade-school maths task was scored two ways: leniently, pulling the final number out of wherever the model put it, and strictly, requiring the exact expected output format. The lenient scores across all fourteen builds sit in a band about eleven points wide. The strict scores range from 9.86 to 37.68 — and the full-precision baseline is 24.64, which several quantised builds beat.
The arithmetic, in other words, is largely intact. What moves is whether the answer comes out in the shape something downstream was expecting. A pipeline that parses model output — and most security tooling parses model output — is exposed to the second number, not the first.
That is one run on one model, and strict-match scoring is brittle by design. But the direction of the finding has independent support. A separate 2026 paper proposes measuring the overlap in correct answers between a model and its quantised variant rather than comparing their scores, and reports:
the base and quantized variants usually have a shift in behavior even when accuracy and perplexity are preserved
Two models can post the same average while disagreeing item by item. An aggregate score is not a promise that you will get the same answers.
Why none of this describes the model you are running
The numbers above are for one specific artifact: Llama-3.1-8B-Instruct, quantised with a particular toolchain at a particular version, evaluated on general benchmarks.
They are not a licence to assume a 33B security fine-tune loses three-tenths of a point at Q4_K_M. Different base, different size, different fine-tuning, different domain, and a task nobody in that study measured. The most that transfers is the shape of the effect. The magnitude does not, and neither does the reassurance.
This is the same error as reading a base model's benchmarks as though they described a fine-tune of it, and it earns the same grade here: real evidence, about a different artifact. That is what tier D means, and it is why a model page states the quantisation as a fact about the file rather than folding it into a score.
Compounding is the version that actually bites. A community build is often a fine-tune of an abliterated model, then quantised — two uncontrolled edits stacked, with the pair almost never evaluated together and the headline numbers belonging to something upstream of both.
Reading a quant label
- The letters are the format, not just the width.
Q4_K_MandQ4_0are both "4-bit" and are not the same thing. - Higher is cheap insurance.
Q8_0was statistically indistinguishable from full precision here while still cutting the file nearly in half. If you have the disk, the argument for going lower is footprint, not quality. - Below four bits, expect real loss — and expect it to land hardest on multi-step reasoning.
- Ask whether anything was measured on this file. For almost every community GGUF the answer is no. That is not a reason to avoid it; it is the difference between a known cost and an assumed one.
Sources. Uygar Kurt, "Which Quantization Should I Use? A Unified Evaluation of llama.cpp Quantization on Llama-3.1-8B-Instruct", 2026 — all figures above are read from its results tables. Baha Rababah, Shahzeb Qamar, Lorenz Sparrenberg, Rafet Sifa, Murat Kantarcioglu, Cuneyt Gurcan Akcora and Carson K. Leung, "The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs", 2026 — the quotation is from its abstract. Retrieved 5 September 2026.