All resources
Topic

llm-safety

2 resources across 2 kinds

Tools

  1. Hereticdual-usehigh-risklicence

    Heretic automatically removes safety alignment ("abliteration") from transformer language models by orthogonalizing components against identified refusal directions, using an optimizer that minimizes KL divergence on benign prompts to preserve general capability.

    Open ↗

Frameworks & agents

  1. A Python library providing standardized, tested implementations of 15+ LLM safety judges (StrongREJECT, LlamaGuard, WildGuard, HarmBench, and others) behind a unified API. Returns a normalized 0-1 harm score (p_harmful) for LLM conversations, supports both local fine-tuned and remote foundation-model judges, and warns when a setup diverges from the original implementation to preserve reproducibility.

    Open ↗