Tools
Tool · llm-red-teamingdual-usehigh-risklicence

Heretic

Heretic automatically removes safety alignment ("abliteration") from transformer language models by orthogonalizing components against identified refusal directions, using an optimizer that minimizes KL divergence on benign prompts to preserve general capability.

Use responsibly

High risk of account lockouts, WAF bans, and terms-of-service violations. Requires explicit authorization and small, targeted inputs.

More tools