All resources
Topic

refusal-direction

1 resources across 1 kinds

Tools

  1. Hereticdual-usehigh-risklicence

    Heretic automatically removes safety alignment ("abliteration") from transformer language models by orthogonalizing components against identified refusal directions, using an optimizer that minimizes KL divergence on benign prompts to preserve general capability.

    Open ↗