All field notes

Guide · Foundations · D. Rose · 31 August 2026 · Updated 5 September 2026 · 4 min

What abliteration actually does to a model

Refusal sits on a single direction in the residual stream. Deleting it removes a behaviour, not a limitation — which is why an uncensored model is more willing, never more capable.

The word turns up on model cards constantly — "abliterated", "refusal-direction de-risk at the weight level", "uncensored build of X". It is used across this catalog too, on a good half-dozen entries. It deserves an actual explanation, because what it does is narrow, specific, and very easy to over-read.

The finding underneath it

In 2024 a group of researchers looked at how refusal is represented inside chat models and found something surprisingly simple. Across thirteen open-source chat models, up to 72B parameters, refusal turned out to be mediated by a one-dimensional subspace — a single direction in the model's residual stream, the running vector that each layer reads from and writes back to.

The paper's own summary of the effect is the part worth holding onto:

we find a single direction such that erasing this direction from the model's residual stream activations prevents it from refusing harmful instructions, while adding this direction elicits refusal on even harmless instructions.

Both halves matter. Erase the direction and the model stops declining. Add it and the model starts declining perfectly innocuous requests. That two-way result is what makes the claim credible: it is not "we found something correlated with refusal", it is "we found the lever, and it moves in both directions".

Abliteration is that finding used in one direction, baked into the weights so it ships in the file. You do not need the original training data, a training run, or a preference dataset. You need the model, a set of prompts it refuses, a set it does not, and the arithmetic to find the direction that separates them.

What it therefore does not do

This is the part the model cards leave out, and it follows directly from the mechanism.

Removing the refusal direction removes a behaviour. It does not add knowledge. It does not add reasoning. It does not teach the model anything about exploitation, or malware, or a protocol it had never seen. Nothing in the procedure introduces new information — it deletes the model's ability to represent one particular thing, and that thing is "decline this".

So an abliterated model is more willing than its parent. It is not more capable than its parent. Those are different properties and only one of them is what you want when you ask a security question and need the answer to be right.

A model that answers every question is not thereby a model that answers correctly. A witness who never says "no comment" has not become a reliable witness.

This is why a compliance rate — "100% compliance, 0% refusal" — is not a capability score, and why the catalog files it separately from the evidence tiers rather than as one of them. It is a real measurement of a real property. The property is willingness.

The caveat that gets transferred and shouldn't

The paper describes its method as one that "surgically disables refusal with minimal effect on other capabilities", and that reassurance gets repeated a great deal in model-card prose.

It should not travel that far. "Minimal effect on other capabilities" is what the authors measured, for the procedure they describe, on the models and evaluations in their paper. A community GGUF whose card says "refusal-direction de-risk at the weight level" is not that procedure, was not run through those evaluations, and in most cases documents neither the prompts used to derive the direction nor any measurement of what happened afterwards.

Reading the paper's result as a guarantee about a particular downloaded file is the same mistake as reading a base model's benchmark scores as though they described a fine-tune of it. Real evidence, about a different artifact. It is the single most common error in this space, it is why this site grades that situation as tier D rather than as evidence, and it applies to reassurance exactly as much as it applies to scores.

Two further things are usually unmeasured in practice, and are worth knowing you don't know:

  • How much else moved. The direction is derived from a specific prompt set. A different set finds a slightly different direction, and what else lies along it is not something a card usually reports.
  • What the quantisation did. Most abliterated models are distributed as 4-bit or lower GGUFs. That is a second uncontrolled edit stacked on the first, and the pair is almost never evaluated together.

How to read a card that says it

Three questions, and they are the same three that work on any model card:

  1. What exactly was edited? A named base, a named procedure, and ideally the prompt sets. "De-risked at the weight level" is a description of nothing.
  2. What was measured afterwards, on this artifact? Not the base model's scores. Not the paper's. This build's.
  3. Is the number a capability number or a compliance number? If the only figure is a refusal rate, then what has been demonstrated is willingness, and the capability question is simply open.

Most cards answer none of the three. That is not a reason to avoid these models — willingness is genuinely useful, and a model that engages with an authorized security question beats one that lectures you about it. It is a reason to be precise about what you are getting: a model that will answer, with the accuracy of its answers still an open question.


Source. Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, Neel Nanda, "Refusal in Language Models Is Mediated by a Single Direction", NeurIPS 2024. Quotations above are from the paper's abstract. Retrieved 31 August 2026.

More in Foundations