Defending against prompt injection: detectors, guards, and the difference

D. Rose · 16 August 2026 · 7 min

The open-weight defensive layer is real and growing, but it is three different jobs that get sold as one. A cited tour of the dedicated injection detectors, the broader content-policy guards, and the evaluation judges that are not runtime filters — plus the attack generators they are tested against, and how the pieces actually stack in a deployment.

Most of the attention on open security models goes to the offensive side. The defensive layer is where a real deployment lives — the models that sit in front of and behind an LLM to catch an attack before it reaches a tool call. That layer is now populated with open weights, and the single most useful thing to know about it is that it is three different jobs, not one, and they are routinely sold as interchangeable. They are not.

Three jobs, not one

  • An injection detector answers a narrow question: is this input trying to hijack the model's instructions? It is usually a small encoder classifier on the input side, and it does not know or care whether the requested content is otherwise harmful.
  • A content-policy guard answers a broader question: does this prompt or response violate a harm taxonomy — violence, CSAM, self-harm, sometimes cyberattack? It is a policy classifier, typically larger, and it is not injection-specific.
  • An evaluation judge answers an offline question: did an attack succeed, scored over a benchmark? It is a scoring model for red-team research, not a filter you put in the request path.

Conflating these is the defensive-side version of quoting a base model's benchmark for a fine-tune. A jailbreak detector will happily pass a well-phrased request for malware because judging the payload is not its job. A content guard will miss a novel injection string that carries no policy-flagged content. And a HarmBench-style judge is built to grade transcripts after the fact, not to sit inline. Pick the wrong one for the slot and you have a guardrail that measures the wrong thing.

What they are tested against: the attack-generator side

The defensive models exist because there is an open, active attack-generator side to push against — the red half of a red/blue loop. These are open-weight models whose entire purpose is to produce adversarial prompts for resilience testing.

Others in this group are narrower or lower-confidence: Knowledge-to-Jailbreak (THU-KEG, arXiv 2406.11682) turns domain knowledge into jailbreak prompts; Jailbreak-R1 is a gated RL attack generator (arXiv 2506.00782); the Qwen3 Injection Generator (Wambo Security) targets secret-extraction evals and is documented by its card only. The S4MPL3BI4S Llama 3 Prompt Injection Generator is community, low-confidence — no dataset, eval, or paper. The point of naming them is not to catalog attacks; it is that the detectors below are only as trustworthy as the adversarial distribution they were trained and measured against, and that distribution is public.

Dedicated injection and jailbreak detectors

These do one job: flag an input as an injection or jailbreak attempt. Most are small encoder classifiers you can run cheaply on every request. Sizes and behavior are from the model cards; none of these should be read as a general content filter.

DetectorWhat it isScopeSource
Meta Prompt Guardv1 mDeBERTa-v3-base 86M (benign/injection/jailbreak, 3-class); v2 mDeBERTa 86M + DeBERTa-v3-xsmall 22M (binary)Direct injection + jailbreak86M, v2-86M, v2-22M
PIGuard / InjecGuardDeBERTa-v3-base ~184M, trained specifically against over-defense false positivesInjectionleolee99/PIGuard, arXiv 2410.22770
Sentinelv1 ModernBERT-large 395M / 8K; v2 a Qwen3-0.6B generative detector / 32K + GGUFInjectionqualifire/prompt-injection-sentinel, arXiv 2506.05446
Protect AI LLM Guard scannersDeBERTa-v3-base v1/v2 + small-v2Direct injection only (card-stated)protectai, github
NeMoGuard JailbreakDetectRandom forest over Snowflake Arctic Embed M Long; ONNXJailbreaknvidia/NemoGuard-JailbreakDetect, arXiv 2412.01547

The rest of the group fills in the size and license spectrum. The small BERT-class baselines are the cheapest option and the least ambitious: jackhhao/jailbreak-classifier is a BERT-base-uncased ~110M classifier, and fmops/distilbert-prompt-injection is a DistilBERT ~67M baseline. deepset's injection detector is another DeBERTa-v3-base ~184M. Patronus AI's Wolf Defender / Lion Warden are mmBERT 0.3B/0.1B/ONNX multi-head threat classifiers; Samsung SDS's SGuard pairs a Granite-3.3-2B JailbreakFilter with a ContentFilter across 60+ patterns (arXiv 2511.12497); DataSentinel (Mistral 7B v0.1) targets indirect injection via a minimax game (arXiv 2504.11358); and DARWIN-Guard is a Qwen3Guard-Gen-8B fine-tune aimed at evolving attacks (arXiv 2607.19829).

Two hedges to carry forward. Access is not uniform. Vijil Dome (ModernBERT-base ~0.1B) is gated, and Semalith v1.5 (DeBERTa-v3-base 184M, 22 classes) is gated and research-only (arXiv 2607.22545) — open-weight in the taxonomy sense is not the same as OSI-open or freely downloadable. Scope is narrower than the name. Protect AI's scanners are card-stated for direct injection only, and DataSentinel is aimed at the indirect case; a detector tuned for one does not automatically cover the other.

Broader safety guards

These are the content-policy classifiers. They judge whether a prompt or response is harmful against a taxonomy, on the input side, the output side, or both. Several include a jailbreak or injection category, but that is one label among many — it does not make them dedicated injection detectors.

GuardBackbone / sizesWhat it coversSource
Meta Llama GuardLlama2-7B v1; Llama3-8B v2; v3 8B/1B/11B-Vision; v4 12B (Llama 4 Scout, pruned)Input/output harm taxonomyarXiv 2312.06674, github
Qwen3GuardGen 0.6B/4B/8B + Stream 0.6B/4B/8B; 119 languagesHarm categories incl. jailbreak inputsarXiv 2510.14276
Granite Guardian (IBM)2B / 3B-A800M / 5B / 8B, v3.0–4.1 + HAP classifiersInput/output harm, jailbreak, RAG, hallucinationarXiv 2412.07724
ShieldGemma (Google)Gemma 2 2B/9B/27B4 harm categoriesarXiv 2407.21772
WildGuard (AllenAI)Mistral 7B v0.3Prompt/response harm + refusal; taxonomy includes cyberattackarXiv 2406.18495

The broader guard field extends well past these five — OpenAI's policy-conditioned gpt-oss-safeguard MoE (20B/120B), NVIDIA's Aegis / NeMoGuard line, and multilingual edge guards like WalledGuard, HaloGuard, PolyGuard and DuoGuard — but the five above are the load-bearing ones for a general deployment. The size range is the deployment lever: a 0.6B Qwen3Guard or a 2B ShieldGemma is cheap enough to run on every turn, while an 8B Llama Guard or Granite Guardian is a heavier check you might reserve for the output side.

The judges are not filters

One category in the taxonomy looks like a guard and is not. HarmBench classifiers (CAIS — Llama2 13B + Mistral 7B) are evaluation judges: they score whether an attack succeeded across a benchmark. They are explicitly not runtime filters. The same is true of the judge family more broadly — ShieldLM (Tsinghua) and MD-Judge (OpenSafetyLab, on SALAD-Bench) are built to grade transcripts, not to gate live traffic.

This matters in both directions. Putting a judge inline gets you a slow, offline-shaped model doing a job it was not calibrated for. And — the more common error — reading a guard's benchmark score as if it were a judge's leaderboard number conflates "how well it filters" with "how well it grades." When a security model's card reports a strong HarmBench figure only when paired with an external Llama Guard, as some do, that is the tell: the base model is not meant to be the sole safety layer, and the judge score and the filter are two different measurements.

How the pieces stack

The honest deployment pattern is layered, because each of the three jobs above covers a gap the others leave open. For a serious deployment, the strongest arrangement is:

  1. A dedicated injection detector on the input side — Prompt Guard 2, PIGuard, or Sentinel — to catch instruction-hijacking before it reaches the model.
  2. A separate content-policy guard — Qwen3Guard, Llama Guard, or Granite Guardian — to check prompts and responses against a harm taxonomy.
  3. Task-specific cyber models kept behind operational controls — tool permissions, sandboxing, and audit logs — so that a model's output cannot act on the world unmediated.

No single model in this survey does all three, and none of them is a substitute for the controls in step three. A detector reduces the rate at which an attack reaches your model; a guard reduces the rate at which harmful content leaves it; permissions, sandboxing and logging bound what happens when both are wrong. That layering — not any one classifier's score — is the defensive posture worth deploying.

This is the half of the ecosystem that the offensive catalog does not show. AdversariaLLM is being built to measure both, and the defensive layer is measured the same way as the offensive one: every model named traces to a card or a paper, the injection detector is never confused with the content guard, and the evaluation judge is never mistaken for a filter.