Guide · Prompt security · D. Rose · 19 July 2026 · Updated 12 August 2026 · 1 min
A checklist for prompt-injection reviews
The eight questions I ask of any system that puts untrusted text in front of a model.
Prompt injection is not one bug; it is a class, and you review for a class with a checklist, not a clever test case.
- Where does untrusted text enter the prompt?
- Can that text reach a system or developer instruction?
- What can the model do — tools, retrieval, code execution?
- Is any model output ever treated as trusted downstream?
- What is the blast radius of the most capable tool?
- Is there a human in the loop before an irreversible action?
- Do you log enough to reconstruct an incident without logging the prompt content itself?
- What happens on the second injected instruction, after the first one was caught?
None of these are exotic. The failures I have seen were all a missing answer to one of them, discovered after the fact.