Framework · llm-red-teaming
JudgeZoo
A Python library providing standardized, tested implementations of 15+ LLM safety judges (StrongREJECT, LlamaGuard, WildGuard, HarmBench, and others) behind a unified API. Returns a normalized 0-1 harm score (p_harmful) for LLM conversations, supports both local fine-tuned and remote foundation-model judges, and warns when a setup diverges from the original implementation to preserve reproducibility.
More frameworks & agents
- PyRITllm-red-teamingPython Risk Identification Tool for generative AI — a framework for automating red-team assessments of LLM systems; named alongside garak.
- Agentic Malware Analysisagent-security-skillsKali-based Docker environment with 50+ RE tools, an MCP-connected Binary Ninja or Ghidra backend, and an orchestrator skill that lets Claude Code or Codex CLI turn a binary into a case directory of ranked evidence and hypotheses.
- Anthropic Cybersecurity Skillsagent-security-skillsApache-2.0 library of 818 agentskills.io-format cybersecurity skills in 34 domains, mapped to MITRE ATT&CK, NIST CSF 2.0, ATLAS, D3FEND, NIST AI RMF and F3, for loading into Claude Code, Codex CLI, Cursor and similar agents.
- AppSec (florianbuetow)agent-security-skillsMIT Claude Code plugin bundling 62 slash-command skills across OWASP, STRIDE, PASTA, LINDDUN, MITRE ATT&CK and CWE Top 25, plus six red-team persona agents, for reviewing a codebase and generating fixes.
- Arm Metiswhite-box-sastApache-2.0 white-box framework with broad language support, local-model support, deterministic evidence collection, and validation of SAST findings.
- Awesome Claude Securityagent-security-skillsClaude Code plugin marketplace of 45 installable plugins (110 skills) covering pentest, threat modeling, detection engineering, DFIR, GRC and LLM/agentic-AI security, with role bundles such as pentester that auto-install their parts.