All resources
Topic

ai-security

13 resources across 3 kinds

Tools

  1. g0 (Guard0)activedual-use

    Open-source AI agent security assessment tool that scans AI/MCP codebases with rule-based static analysis (A–F grading) and runs 3,900+ adversarial payloads for red-team testing against live agents. Also does MCP supply-chain discovery, AI-BOM generation (CycloneDX), and proxy-based runtime enforcement.

    Open ↗
  2. Claude Code skill and CLI scanner that audits community skill files across 26 detection categories (prompt injection, exfiltration, permission bypass), with runtime hooks that block dangerous commands and inspect MCP responses.

    Open ↗
  3. is-maliciouscloud costpassive

    MIT-licensed Node CLI that sends a project's selected source, config, build, and CI files to TypeSafe Jev and reports suspicious files and line ranges with probabilities; needs a TypeSafe API key and spends paid input tokens.

    Open ↗

Frameworks & agents

  1. Claude Code plugin marketplace of 45 installable plugins (110 skills) covering pentest, threat modeling, detection engineering, DFIR, GRC and LLM/agentic-AI security, with role bundles such as pentester that auto-install their parts.

    Open ↗
  2. Pack of 38 defensive-only Markdown skills for Claude Code and compatible coding agents, covering MCP security, prompt-injection defense, OWASP LLM Top 10, VPS/WordPress/Cloudflare hardening and incident response.

    Open ↗
  3. Settings template, blocking hooks, path-scoped language rules, slash commands and an MCP template for Claude Code, with notes on sandboxing bypass-permissions runs, local models and context management.

    Open ↗
  4. Jev Use Casescloud cost

    MIT-licensed Python package of 37 typed-decision runners for the paid TypeSafe Jev API, including seven SOC agents that recommend triage, containment and escalation for the caller to run, and three that screen prompts and drafts.

    Open ↗

References

  1. A curated list of AI/ML security resources organized into adversarial examples, evasion attacks, poisoning attacks, feature selection, and code (e.g. CleverHans, Foolbox). Aggregates papers, libraries, videos, and blog posts on attacking and defending ML systems.

    Open ↗
  2. An empirical, practical guide to LLM hacking from security firm Forces Unseen, covering prompt injection, offensive and defensive techniques, and interactive playgrounds. Source for the handbook hosted at doublespeak.chat.

    Open ↗
  3. Dataset of 15,140 in-the-wild prompts collected from Reddit, Discord, websites and open-source datasets, 1,405 of them jailbreaks, plus a 390-question forbidden-scenario set for measuring jailbreak effectiveness.

    Open ↗
  4. Jailbreaking Frontier Modelsactivecloud costdual-usehigh-risklicence

    Dataset of harmful-behaviour prompts (drug, chemical, biological, radiological, nuclear, explosive) and a reference PRBO reward function for training jailbreaking agents; the RL training loop is not included.

    Open ↗
  5. README-only collection of copy-paste jailbreak prompts for DeepSeek R1, Grok 3, Gemini 2.0, ChatGPT (DAN), Claude 2 and Llama 2, plus a Gemini system-prompt leak prompt, mostly reposted from linked Reddit, blog and GitHub posts.

    Open ↗
  6. Vendor-organised collection of files, mostly Markdown, presenting what it says are the verbatim system prompts of major chatbots and coding agents (Anthropic, OpenAI, Google, xAI, Cursor, Kimi and others), with a dated additions table and an open invitation to PRs.

    Open ↗