Data, Pre-Training & Post-Training · 7 min

Shrinking Weights for Inference: GPTQ, AWQ, and GGUF

Post-training quantization drops weights to 4-bit for a few percent of quality, and GGUF is a shipping container, not a compression method.

A 7B model in FP16 is about 14 GB of weights. The same model at 4-bit is roughly 4.4 GB. That gap is the whole reason quantization exists: the weights are the thing that won't fit, and inference is memory-bandwidth bound long before it is compute bound. Every parameter you pull from VRAM to compute one token is a parameter you paid for in bytes moved. Halve the bytes and you roughly halve the read traffic, so the model gets smaller and faster. That is not the usual engineering trade, where you buy one with the other.

The catch is that a weight learned as a 16-bit float now has to be one of only 16 values (INT4). You are throwing information away. The craft of post-training quantization (PTQ) is deciding which information to throw away so the model barely notices.

The mental model: rounding is a budget

Start with the naive scheme. Take a block of weights, find the largest magnitude, and map the range linearly onto the integer grid:

scale = max(|w|) / 7            # signed INT4 spans -8..7
q     = round(w / scale)        # store the int
w'    = q * scale               # dequantize at inference

This is round-to-nearest (RTN). It works well down to 8-bit and falls apart at 4-bit, because a few weights matter enormously and RTN spends identical precision on all of them. GPTQ and AWQ are two different answers to the same question: how do you spend precision where it matters?

GPTQ: round one weight, then fix the rest

GPTQ descends from Optimal Brain Quantization. The insight is that quantizing a weight introduces a known error, and you can compensate by nudging the weights you have not quantized yet. It runs layer by layer, minimizing the squared difference between the original layer output and the quantized one over a small calibration set, typically 128 sequences of real text.

for each column j (left to right):
    quantize w_j  -> nearest grid point
    err = w_j - w_j_quantized
    push a correction proportional to err into columns j+1..n
    (correction scaled by the inverse Hessian of the layer inputs)

The Hessian here is XᵀX, the covariance of the layer's inputs over the calibration data. It encodes which input directions the layer actually sees, so the correction knows which surviving weights can absorb the error most cheaply. The original OBQ did this in an optimal per-row order. GPTQ's practical trick was noticing that for large models a fixed left-to-right order works just as well, which lets every row reuse one Hessian and one Cholesky factorization. That is what turned a days-long procedure into minutes.

The weakness is visible in the diagram: error propagates forward down the columns. Round a column badly early and the residual chases you through the rest of the layer. On JarvisLabs' Llama-3.1-8B benchmark, 4-bit GPTQ lands at 6.90 WikiText perplexity against 6.56 for FP16, and it is the most brittle of the group on generation: 46.3% pass@1 on HumanEval where FP16 scores 56.1%.

AWQ: protect the salient channels, leave the rest

AWQ starts from a sharper observation. Not all weights matter equally, and you can find the important ones without looking at the weights at all. Look at the activations instead. The weight channels multiplied by consistently large activations are the ones whose errors blow up in the output. Only about 1% of channels are salient in this sense, and protecting those recovers most of the lost accuracy.

You cannot literally keep 1% of channels in FP16 without wrecking your kernel with mixed-precision bookkeeping. AWQ's move is algebraic. Scale the salient weight channels up by a per-channel factor s before rounding, so they land on the grid more precisely, and scale the matching activations down by 1/s at runtime so the math is unchanged: (s·w)·(x/s) = w·x. The scale comes from activation statistics collected offline on a calibration set.

Because AWQ does no backprop and no forward error propagation, just a closed-form rescale and round, the authors show it does not overfit the calibration set the way a reconstruction method can, and it generalizes across domains and modalities. On the same benchmark it comes in at 6.84 perplexity and 51.8% HumanEval, ahead of GPTQ on both. That reliability is why it has become a common default for GPU serving.

Researcher: Calibration data is a real experimental variable. GPTQ's forward correction can overfit 128 samples of English Wikipedia and then underperform on code or other languages. If you quantize your own weights, match calibration data to the deployment distribution and re-measure on a held-out set. Do not trust the single perplexity number on a model card.

GGUF is a box, not a method

Here is the distinction the lesson turns on. GPTQ and AWQ are algorithms; they decide the integer values. GGUF is a file format. It answers "how is this shipped," not "how were the weights compressed."

A .gguf file is a self-contained container: a header, a block of key-value metadata (including the tokenizer and the chat template), a tensor index, and the quantized weight blobs, aligned to 32 bytes so the runtime can mmap them and page in only the layers it needs.

GGUF FILE (the container)
├─ magic + version
├─ metadata KV   → tokenizer, chat template, hyperparams
├─ tensor index  → name, shape, dtype, offset per tensor
└─ weight blobs  → each tensor quantized with SOME scheme
                    e.g. Q4_K, Q6_K, F16  (mixed per tensor)

That "SOME scheme" is where people conflate things. The compression inside a GGUF is usually a k-quant, and that is the algorithm. K-quants use two-level block scaling. Weights are grouped into a super-block of 256 carrying one FP16 scale and one FP16 minimum, split into 8 sub-blocks of 32 that each carry a 6-bit sub-scale. A weight is reconstructed roughly as super_min + super_scale · (sub_min + sub_scale · code). For Q4_K the codes are 4-bit, and once you count the scale overhead the effective rate is about 4.5 bits per weight, not 4.

The suffix decodes cleanly. Q4 is 4-bit codes, K is the k-quant two-level scheme, _M is the "medium" mixed-precision policy. Q4_K_M is not uniformly 4-bit: llama.cpp keeps the value and output projections at Q6_K and the embeddings and norms higher, spending bits where the model is fragile. That hand-tuned allocation pushes the real rate to about 4.8–4.9 bits per weight (roughly 4.4 GB for a 7B) while holding quality better than a uniform 4-bit pass. On the JarvisLabs run Q4_K_M posts 6.74 perplexity, the best of the 4-bit methods there.

Builder: "Convert to GGUF" and "quantize to Q4_K_M" are two steps, not one. You can package FP16 weights in a GGUF and compress nothing. GPTQ and AWQ ship as safetensors for vLLM or TGI on GPUs; GGUF with k-quants is the format for llama.cpp on CPU, Metal, and consumer GPUs. Pick the format for the runtime, then pick the bit-width.

What actually goes wrong

Safety is not invariant under quantization. Alignment behavior lives in the same weights, and dropping to 4-bit can erode refusals and jailbreak resistance even when perplexity looks fine. A model that is "the same quality" on your eval can be measurably more compliant to attacks.

Defender: Re-run your refusal and jailbreak suite against the quantized artifact you will deploy, not the FP16 checkpoint. Treat each bit-width as a separate model for red-teaming.

GGUF files are untrusted input. The container carries a chat template and tokenizer that your runtime parses, and malformed GGUF headers have produced real heap-overflow bugs in ggml and llama.cpp. A poisoned chat template injects instructions into every prompt. A .gguf from a random uploader is code-adjacent data: verify the source and pin the hash.

None of these methods is universally best. AWQ generally edges GPTQ on modern instruction-tuned models for GPU serving; k-quants own local inference. The honest headline is that all of them stay within a few percent of FP16, which is why quantized models, not the FP16 originals, are what most people actually run. How those weights move through the KV cache and sampler at decode time is the next module's problem.

Sources