red-teaming
2 resources across 1 kinds
Tools
- Open ↗LLMartdual-use
Intel Labs' LLM adversarial-robustness toolkit built on PyTorch/Hugging Face, implementing discrete-optimization attacks (notably GCG) and soft-prompt/adversarial-suffix optimization to red-team text LLMs, VLMs, and diffusion models at scale, with AdvBench/HarmBench dataset integrations and CLI/programmatic interfaces.
- Open ↗Hereticdual-usehigh-risklicence
Heretic automatically removes safety alignment ("abliteration") from transformer language models by orthogonalizing components against identified refusal directions, using an optimizer that minimizes KL divergence on benign prompts to preserve general capability.