All resources
Topic

jailbreak

2 resources across 2 kinds

Tools

  1. promptmap2activedual-use

    promptmap2 is an automated vulnerability scanner for custom LLM applications that runs prompt-injection attacks against a target (white-box with known system prompt, or black-box HTTP endpoint). It uses a dual-LLM setup where a controller LLM judges whether attacks succeeded, with 50+ pre-built rules across categories like prompt stealing, jailbreak, and harmful-content generation; supports OpenAI, Anthropic, Gemini, Grok, and Ollama models.

    Open ↗

References

  1. Primary research from the UK AI Security Institute (DSIT): frontier-model cyber-capability evaluations, safeguard and jailbreak testing, and incident reports from their own agent testing.

    Open ↗