Tool · llm-red-teamingdual-usehigh-risklicence
Heretic
Heretic automatically removes safety alignment ("abliteration") from transformer language models by orthogonalizing components against identified refusal directions, using an optimizer that minimizes KL divergence on benign prompts to preserve general capability.
Use responsibly
High risk of account lockouts, WAF bans, and terms-of-service violations. Requires explicit authorization and small, targeted inputs.
More tools
- g0 (Guard0)llm-red-teamingOpen-source AI agent security assessment tool that scans AI/MCP codebases with rule-based static analysis (A–F grading) and runs 3,900+ adversarial payloads for red-team testing against live agents. Also does MCP supply-chain discovery, AI-BOM generation (CycloneDX), and proxy-based runtime enforcement.
- garakllm-red-teamingLLM vulnerability / red-team scanner (standalone CLI) for probing generative-AI models for jailbreaks, data leakage, and other failure modes.
- Kimi Thinking Prefillllm-red-teamingClient-side SillyTavern extension that rewrites outgoing Kimi/Moonshot chat requests to prefill an assistant reasoning_content turn, so the model continues thinking from user-seeded reasoning without a server patch.
- LLMartllm-red-teamingIntel Labs' LLM adversarial-robustness toolkit built on PyTorch/Hugging Face, implementing discrete-optimization attacks (notably GCG) and soft-prompt/adversarial-suffix optimization to red-team text LLMs, VLMs, and diffusion models at scale, with AdvBench/HarmBench dataset integrations and CLI/programmatic interfaces.
- promptmap2llm-red-teamingpromptmap2 is an automated vulnerability scanner for custom LLM applications that runs prompt-injection attacks against a target (white-box with known system prompt, or black-box HTTP endpoint). It uses a dual-LLM setup where a controller LLM judges whether attacks succeeded, with 50+ pre-built rules across categories like prompt stealing, jailbreak, and harmful-content generation; supports OpenAI, Anthropic, Gemini, Grok, and Ollama models.
- agent-chaperoneagent-guardrailApache-2.0 MCP proxy and Claude Code hooks adapter that screens an agent's tool calls before they run and tool results before it reads them against thresholds in a policy file; log-only by default, needs a TypeSafe API key.