All field notes

Guide · Prompt security · D. Rose · 19 July 2026 · Updated 12 August 2026 · 1 min

A checklist for prompt-injection reviews

The eight questions I ask of any system that puts untrusted text in front of a model.

Prompt injection is not one bug; it is a class, and you review for a class with a checklist, not a clever test case.

  1. Where does untrusted text enter the prompt?
  2. Can that text reach a system or developer instruction?
  3. What can the model do — tools, retrieval, code execution?
  4. Is any model output ever treated as trusted downstream?
  5. What is the blast radius of the most capable tool?
  6. Is there a human in the loop before an irreversible action?
  7. Do you log enough to reconstruct an incident without logging the prompt content itself?
  8. What happens on the second injected instruction, after the first one was caught?

None of these are exotic. The failures I have seen were all a missing answer to one of them, discovered after the fact.

More in Prompt security