All resources
Topic

prompt-injection

12 resources across 4 kinds

Tools

  1. promptmap2activedual-use

    promptmap2 is an automated vulnerability scanner for custom LLM applications that runs prompt-injection attacks against a target (white-box with known system prompt, or black-box HTTP endpoint). It uses a dual-LLM setup where a controller LLM judges whether attacks succeeded, with 50+ pre-built rules across categories like prompt stealing, jailbreak, and harmful-content generation; supports OpenAI, Anthropic, Gemini, Grok, and Ollama models.

    Open ↗
  2. Claude Code skill and CLI scanner that audits community skill files across 26 detection categories (prompt injection, exfiltration, permission bypass), with runtime hooks that block dangerous commands and inspect MCP responses.

    Open ↗
  3. agent-chaperonecloud costhosted

    Apache-2.0 MCP proxy and Claude Code hooks adapter that screens an agent's tool calls before they run and tool results before it reads them against thresholds in a policy file; log-only by default, needs a TypeSafe API key.

    Open ↗
  4. dsh-jev-interceptorcloud costhostedopt-in

    MIT-licensed DeepSeek Harness plugin that scores non-read-only tool calls for risk, irreversibility and injection suspicion with TypeSafe's hosted Jev model, escalating or denying only in enforce mode; needs a paid API key.

    Open ↗
  5. Jev Sentinelcloud costhosted

    MIT-licensed guard for coding agents (a Pi extension and Claude Code/Codex CLI plugins) that scores tool calls, outputs and replies with TypeSafe's Jev for injection and risk, then allows, asks, or blocks; needs a TypeSafe key.

    Open ↗
  6. jev-edgecloud cost

    Apache-2.0 fail-open gateway filter for OpenResty, APISIX, Kong, Envoy and JS hosts that scores inbound LLM requests for prompt injection, judging with a paid TypeSafe Jev key, a fine-tuned Laya or an OpenAI-compatible endpoint.

    Open ↗
  7. jev-guardcloud costhosted

    MIT-licensed hook for seven coding agents and any ACP pair that asks TypeSafe's hosted Jev typed questions about a tool call or its result, then denies, asks or allows it and flags suspected prompt injection; needs a billed key.

    Open ↗

Frameworks & agents

  1. jev-harnesscloud costopt-in

    MIT-licensed research-stage harness in which an LLM proposes one read or patch, code validates it, Jev answers four yes/no questions, and a fixed table decides; the host owns authorization, and nothing is applied or executed.

    Open ↗

Benchmarks

  1. MIT-licensed one-notebook benchmark that sends 116 labeled prompts to TypeSafe's hosted system_one API and reports per-call latency plus a confusion matrix of its prompt-injection verdicts; needs a TypeSafe account.

    Open ↗
  2. jev-sec-benchcloud costhosted

    MIT-licensed Go benchmark that scores TypeSafe's hosted Jev model on 662 labelled messages (263 injections) and 200 matched vulnerable/secure code pairs; needs a TypeSafe API key, and the documented runner is not in the tree.

    Open ↗

References

  1. An empirical, practical guide to LLM hacking from security firm Forces Unseen, covering prompt injection, offensive and defensive techniques, and interactive playgrounds. Source for the handbook hosted at doublespeak.chat.

    Open ↗
  2. A research paper proposing a component model for the structured analysis of prompt-injection attacks — decomposing an injection into its functional parts (the framing that terminates trusted context, fake system tags, the payload) so defenders can reason about and detect them systematically.

    Open ↗