Tool · llm-red-teamingactivedual-use
g0 (Guard0)
Open-source AI agent security assessment tool that scans AI/MCP codebases with rule-based static analysis (A–F grading) and runs 3,900+ adversarial payloads for red-team testing against live agents. Also does MCP supply-chain discovery, AI-BOM generation (CycloneDX), and proxy-based runtime enforcement.
Use responsibly
Test only systems you own or are explicitly authorized to test. Unauthorized testing is illegal.
Dual-use security content. We link to the upstream project rather than re-hosting ready-to-fire files. Use only in authorized security testing.
More tools
- garakllm-red-teamingLLM vulnerability / red-team scanner (standalone CLI) for probing generative-AI models for jailbreaks, data leakage, and other failure modes.
- Hereticllm-red-teamingHeretic automatically removes safety alignment ("abliteration") from transformer language models by orthogonalizing components against identified refusal directions, using an optimizer that minimizes KL divergence on benign prompts to preserve general capability.
- Kimi Thinking Prefillllm-red-teamingClient-side SillyTavern extension that rewrites outgoing Kimi/Moonshot chat requests to prefill an assistant reasoning_content turn, so the model continues thinking from user-seeded reasoning without a server patch.
- LLMartllm-red-teamingIntel Labs' LLM adversarial-robustness toolkit built on PyTorch/Hugging Face, implementing discrete-optimization attacks (notably GCG) and soft-prompt/adversarial-suffix optimization to red-team text LLMs, VLMs, and diffusion models at scale, with AdvBench/HarmBench dataset integrations and CLI/programmatic interfaces.
- promptmap2llm-red-teamingpromptmap2 is an automated vulnerability scanner for custom LLM applications that runs prompt-injection attacks against a target (white-box with known system prompt, or black-box HTTP endpoint). It uses a dual-LLM setup where a controller LLM judges whether attacks succeeded, with 50+ pre-built rules across categories like prompt stealing, jailbreak, and harmful-content generation; supports OpenAI, Anthropic, Gemini, Grok, and Ollama models.
- agent-chaperoneagent-guardrailApache-2.0 MCP proxy and Claude Code hooks adapter that screens an agent's tool calls before they run and tool results before it reads them against thresholds in a policy file; log-only by default, needs a TypeSafe API key.