jailbreak
2 resources across 2 kinds
Tools
- Open ↗promptmap2activedual-use
promptmap2 is an automated vulnerability scanner for custom LLM applications that runs prompt-injection attacks against a target (white-box with known system prompt, or black-box HTTP endpoint). It uses a dual-LLM setup where a controller LLM judges whether attacks succeeded, with 50+ pre-built rules across categories like prompt stealing, jailbreak, and harmful-content generation; supports OpenAI, Anthropic, Gemini, Grok, and Ollama models.
References
- Open ↗
Primary research from the UK AI Security Institute (DSIT): frontier-model cyber-capability evaluations, safeguard and jailbreak testing, and incident reports from their own agent testing.