Tool · llm-red-teaminglicence
Kimi Thinking Prefill
Client-side SillyTavern extension that rewrites outgoing Kimi/Moonshot chat requests to prefill an assistant reasoning_content turn, so the model continues thinking from user-seeded reasoning without a server patch.
More tools
- g0 (Guard0)llm-red-teamingOpen-source AI agent security assessment tool that scans AI/MCP codebases with rule-based static analysis (A–F grading) and runs 3,900+ adversarial payloads for red-team testing against live agents. Also does MCP supply-chain discovery, AI-BOM generation (CycloneDX), and proxy-based runtime enforcement.
- garakllm-red-teamingLLM vulnerability / red-team scanner (standalone CLI) for probing generative-AI models for jailbreaks, data leakage, and other failure modes.
- Hereticllm-red-teamingHeretic automatically removes safety alignment ("abliteration") from transformer language models by orthogonalizing components against identified refusal directions, using an optimizer that minimizes KL divergence on benign prompts to preserve general capability.
- LLMartllm-red-teamingIntel Labs' LLM adversarial-robustness toolkit built on PyTorch/Hugging Face, implementing discrete-optimization attacks (notably GCG) and soft-prompt/adversarial-suffix optimization to red-team text LLMs, VLMs, and diffusion models at scale, with AdvBench/HarmBench dataset integrations and CLI/programmatic interfaces.
- promptmap2llm-red-teamingpromptmap2 is an automated vulnerability scanner for custom LLM applications that runs prompt-injection attacks against a target (white-box with known system prompt, or black-box HTTP endpoint). It uses a dual-LLM setup where a controller LLM judges whether attacks succeeded, with 50+ pre-built rules across categories like prompt stealing, jailbreak, and harmful-content generation; supports OpenAI, Anthropic, Gemini, Grok, and Ollama models.
- agent-chaperoneagent-guardrailApache-2.0 MCP proxy and Claude Code hooks adapter that screens an agent's tool calls before they run and tool results before it reads them against thresholds in a policy file; log-only by default, needs a TypeSafe API key.