llm-red-teaming
8 resources across 2 kinds
Tools
- Open ↗LLMartdual-use
Intel Labs' LLM adversarial-robustness toolkit built on PyTorch/Hugging Face, implementing discrete-optimization attacks (notably GCG) and soft-prompt/adversarial-suffix optimization to red-team text LLMs, VLMs, and diffusion models at scale, with AdvBench/HarmBench dataset integrations and CLI/programmatic interfaces.
- Open ↗g0 (Guard0)activedual-use
Open-source AI agent security assessment tool that scans AI/MCP codebases with rule-based static analysis (A–F grading) and runs 3,900+ adversarial payloads for red-team testing against live agents. Also does MCP supply-chain discovery, AI-BOM generation (CycloneDX), and proxy-based runtime enforcement.
- Open ↗Hereticdual-usehigh-risklicence
Heretic automatically removes safety alignment ("abliteration") from transformer language models by orthogonalizing components against identified refusal directions, using an optimizer that minimizes KL divergence on benign prompts to preserve general capability.
- Open ↗promptmap2activedual-use
promptmap2 is an automated vulnerability scanner for custom LLM applications that runs prompt-injection attacks against a target (white-box with known system prompt, or black-box HTTP endpoint). It uses a dual-LLM setup where a controller LLM judges whether attacks succeeded, with 50+ pre-built rules across categories like prompt stealing, jailbreak, and harmful-content generation; supports OpenAI, Anthropic, Gemini, Grok, and Ollama models.
- Open ↗Kimi Thinking Prefilllicence
Client-side SillyTavern extension that rewrites outgoing Kimi/Moonshot chat requests to prefill an assistant reasoning_content turn, so the model continues thinking from user-seeded reasoning without a server patch.
Frameworks & agents
- Open ↗
A Python library providing standardized, tested implementations of 15+ LLM safety judges (StrongREJECT, LlamaGuard, WildGuard, HarmBench, and others) behind a unified API. Returns a normalized 0-1 harm score (p_harmful) for LLM conversations, supports both local fine-tuned and remote foundation-model judges, and warns when a setup diverges from the original implementation to preserve reproducibility.