All resources
Topic

llm-red-teaming

8 resources across 2 kinds

Tools

  1. LLM vulnerability / red-team scanner (standalone CLI) for probing generative-AI models for jailbreaks, data leakage, and other failure modes.

    Open ↗
  2. LLMartdual-use

    Intel Labs' LLM adversarial-robustness toolkit built on PyTorch/Hugging Face, implementing discrete-optimization attacks (notably GCG) and soft-prompt/adversarial-suffix optimization to red-team text LLMs, VLMs, and diffusion models at scale, with AdvBench/HarmBench dataset integrations and CLI/programmatic interfaces.

    Open ↗
  3. g0 (Guard0)activedual-use

    Open-source AI agent security assessment tool that scans AI/MCP codebases with rule-based static analysis (A–F grading) and runs 3,900+ adversarial payloads for red-team testing against live agents. Also does MCP supply-chain discovery, AI-BOM generation (CycloneDX), and proxy-based runtime enforcement.

    Open ↗
  4. Hereticdual-usehigh-risklicence

    Heretic automatically removes safety alignment ("abliteration") from transformer language models by orthogonalizing components against identified refusal directions, using an optimizer that minimizes KL divergence on benign prompts to preserve general capability.

    Open ↗
  5. promptmap2activedual-use

    promptmap2 is an automated vulnerability scanner for custom LLM applications that runs prompt-injection attacks against a target (white-box with known system prompt, or black-box HTTP endpoint). It uses a dual-LLM setup where a controller LLM judges whether attacks succeeded, with 50+ pre-built rules across categories like prompt stealing, jailbreak, and harmful-content generation; supports OpenAI, Anthropic, Gemini, Grok, and Ollama models.

    Open ↗
  6. Client-side SillyTavern extension that rewrites outgoing Kimi/Moonshot chat requests to prefill an assistant reasoning_content turn, so the model continues thinking from user-seeded reasoning without a server patch.

    Open ↗

Frameworks & agents

  1. Python Risk Identification Tool for generative AI — a framework for automating red-team assessments of LLM systems; named alongside garak.

    Open ↗
  2. A Python library providing standardized, tested implementations of 15+ LLM safety judges (StrongREJECT, LlamaGuard, WildGuard, HarmBench, and others) behind a unified API. Returns a normalized 0-1 harm score (p_harmful) for LLM conversations, supports both local fine-tuned and remote foundation-model judges, and warns when a setup diverges from the original implementation to preserve reproducibility.

    Open ↗