llm-security
2 resources across 2 kinds
Tools
- Open ↗promptmap2activedual-use
promptmap2 is an automated vulnerability scanner for custom LLM applications that runs prompt-injection attacks against a target (white-box with known system prompt, or black-box HTTP endpoint). It uses a dual-LLM setup where a controller LLM judges whether attacks succeeded, with 50+ pre-built rules across categories like prompt stealing, jailbreak, and harmful-content generation; supports OpenAI, Anthropic, Gemini, Grok, and Ollama models.
References
- Open ↗
A research paper proposing a component model for the structured analysis of prompt-injection attacks — decomposing an injection into its functional parts (the framing that terminates trusted context, fake system tags, the payload) so defenders can reason about and detect them systematically.