# Defending against prompt injection: detectors, guards, and the difference

*D. Rose · 16 August 2026 · 7 min*

> The open-weight defensive layer is real and growing, but it is three different jobs that get sold as one. A cited tour of the dedicated injection detectors, the broader content-policy guards, and the evaluation judges that are not runtime filters — plus the attack generators they are tested against, and how the pieces actually stack in a deployment.

Most of the attention on open security models goes to the offensive side. The defensive layer is where a real deployment lives — the models that sit in front of and behind an LLM to catch an attack before it reaches a tool call. That layer is now populated with open weights, and the single most useful thing to know about it is that it is three different jobs, not one, and they are routinely sold as interchangeable. They are not.

## Three jobs, not one

- **An injection detector** answers a narrow question: is this input trying to hijack the model's instructions? It is usually a small encoder classifier on the *input* side, and it does not know or care whether the requested content is otherwise harmful.
- **A content-policy guard** answers a broader question: does this prompt or response violate a harm taxonomy — violence, CSAM, self-harm, sometimes cyberattack? It is a policy classifier, typically larger, and it is not injection-specific.
- **An evaluation judge** answers an offline question: did an attack succeed, scored over a benchmark? It is a scoring model for red-team research, not a filter you put in the request path.

Conflating these is the defensive-side version of quoting a base model's benchmark for a fine-tune. A jailbreak detector will happily pass a well-phrased request for malware because judging the payload is not its job. A content guard will miss a novel injection string that carries no policy-flagged content. And a HarmBench-style judge is built to grade transcripts after the fact, not to sit inline. Pick the wrong one for the slot and you have a guardrail that measures the wrong thing.

## What they are tested against: the attack-generator side

The defensive models exist because there is an open, active attack-generator side to push against — the red half of a red/blue loop. These are open-weight models whose entire purpose is to produce adversarial prompts for resilience testing.

- **[AmpleGCG / AmpleGCG-Plus](https://huggingface.co/osunlp)** (OSU NLP) are gated Llama-2-7B generators that emit adversarial suffixes via GCG ([arXiv 2404.07921](https://arxiv.org/abs/2404.07921), [2410.22143](https://arxiv.org/abs/2410.22143)). Gated access, not an open download.
- **[ADV-LLM](https://huggingface.co/cesun)** (UCSD/Microsoft) is a set of six generators — Llama 3 8B, Llama 2 7B, Mistral 7B, Guanaco, Vicuna, Phi-3 Mini — that iteratively self-tune jailbreak suffixes ([arXiv 2410.18469](https://arxiv.org/abs/2410.18469), NAACL 2025.naacl-long.297).
- **[Self-RedTeam](https://huggingface.co/mickelliu)** (Qwen2.5-Instruct 3B/7B/14B) runs a self-play attacker/defender loop ([arXiv 2506.07468](https://arxiv.org/abs/2506.07468)).

Others in this group are narrower or lower-confidence: [Knowledge-to-Jailbreak](https://huggingface.co/tsq2000/Jailbreak-generator) (THU-KEG, [arXiv 2406.11682](https://arxiv.org/abs/2406.11682)) turns domain knowledge into jailbreak prompts; [Jailbreak-R1](https://huggingface.co/yukiyounai/Jailbreak-R1) is a gated RL attack generator ([arXiv 2506.00782](https://arxiv.org/abs/2506.00782)); the [Qwen3 Injection Generator](https://huggingface.co/wambosec) (Wambo Security) targets secret-extraction evals and is documented by its card only. The [S4MPL3BI4S Llama 3 Prompt Injection Generator](https://huggingface.co/) is community, low-confidence — no dataset, eval, or paper. The point of naming them is not to catalog attacks; it is that the detectors below are only as trustworthy as the adversarial distribution they were trained and measured against, and that distribution is public.

## Dedicated injection and jailbreak detectors

These do one job: flag an input as an injection or jailbreak attempt. Most are small encoder classifiers you can run cheaply on every request. Sizes and behavior are from the model cards; none of these should be read as a general content filter.

| Detector | What it is | Scope | Source |
|---|---|---|---|
| **Meta Prompt Guard** | v1 mDeBERTa-v3-base 86M (benign/injection/jailbreak, 3-class); v2 mDeBERTa 86M + DeBERTa-v3-xsmall 22M (binary) | Direct injection + jailbreak | [86M](https://huggingface.co/meta-llama/Prompt-Guard-86M), [v2-86M](https://huggingface.co/meta-llama/Llama-Prompt-Guard-2-86M), [v2-22M](https://huggingface.co/meta-llama/Llama-Prompt-Guard-2-22M) |
| **PIGuard / InjecGuard** | DeBERTa-v3-base ~184M, trained specifically against over-defense false positives | Injection | [leolee99/PIGuard](https://huggingface.co/leolee99/PIGuard), [arXiv 2410.22770](https://arxiv.org/abs/2410.22770) |
| **Sentinel** | v1 ModernBERT-large 395M / 8K; v2 a Qwen3-0.6B *generative* detector / 32K + GGUF | Injection | [qualifire/prompt-injection-sentinel](https://huggingface.co/qualifire/prompt-injection-sentinel), [arXiv 2506.05446](https://arxiv.org/abs/2506.05446) |
| **Protect AI LLM Guard scanners** | DeBERTa-v3-base v1/v2 + small-v2 | Direct injection only (card-stated) | [protectai](https://huggingface.co/protectai/deberta-v3-base-prompt-injection), [github](https://github.com/protectai/llm-guard) |
| **NeMoGuard JailbreakDetect** | Random forest over Snowflake Arctic Embed M Long; ONNX | Jailbreak | [nvidia/NemoGuard-JailbreakDetect](https://huggingface.co/nvidia/NemoGuard-JailbreakDetect), [arXiv 2412.01547](https://arxiv.org/abs/2412.01547) |

The rest of the group fills in the size and license spectrum. The small BERT-class baselines are the cheapest option and the least ambitious: [jackhhao/jailbreak-classifier](https://huggingface.co/jackhhao/jailbreak-classifier) is a BERT-base-uncased ~110M classifier, and [fmops/distilbert-prompt-injection](https://huggingface.co/fmops/distilbert-prompt-injection) is a DistilBERT ~67M baseline. [deepset's injection detector](https://huggingface.co/deepset/deberta-v3-base-injection) is another DeBERTa-v3-base ~184M. Patronus AI's [Wolf Defender / Lion Warden](https://huggingface.co/patronus-studio) are mmBERT 0.3B/0.1B/ONNX multi-head threat classifiers; Samsung SDS's [SGuard](https://huggingface.co/SamsungSDS-Research) pairs a Granite-3.3-2B JailbreakFilter with a ContentFilter across 60+ patterns ([arXiv 2511.12497](https://arxiv.org/abs/2511.12497)); [DataSentinel](https://github.com/liu00222/Open-Prompt-Injection) (Mistral 7B v0.1) targets *indirect* injection via a minimax game ([arXiv 2504.11358](https://arxiv.org/abs/2504.11358)); and [DARWIN-Guard](https://huggingface.co/ZJUlilan/DARWIN-Guard) is a Qwen3Guard-Gen-8B fine-tune aimed at evolving attacks ([arXiv 2607.19829](https://arxiv.org/abs/2607.19829)).

Two hedges to carry forward. **Access is not uniform.** [Vijil Dome](https://huggingface.co/vijil/vijil_dome_prompt_injection_detection) (ModernBERT-base ~0.1B) is gated, and [Semalith v1.5](https://huggingface.co/Tejasvi-addagada/semalith-v1.5) (DeBERTa-v3-base 184M, 22 classes) is gated *and* research-only ([arXiv 2607.22545](https://arxiv.org/abs/2607.22545)) — open-weight in the taxonomy sense is not the same as OSI-open or freely downloadable. **Scope is narrower than the name.** Protect AI's scanners are card-stated for *direct* injection only, and DataSentinel is aimed at the *indirect* case; a detector tuned for one does not automatically cover the other.

## Broader safety guards

These are the content-policy classifiers. They judge whether a prompt or response is harmful against a taxonomy, on the input side, the output side, or both. Several include a jailbreak or injection category, but that is one label among many — it does not make them dedicated injection detectors.

| Guard | Backbone / sizes | What it covers | Source |
|---|---|---|---|
| **Meta Llama Guard** | Llama2-7B v1; Llama3-8B v2; v3 8B/1B/11B-Vision; v4 12B (Llama 4 Scout, pruned) | Input/output harm taxonomy | [arXiv 2312.06674](https://arxiv.org/abs/2312.06674), [github](https://github.com/meta-llama/PurpleLlama) |
| **Qwen3Guard** | Gen 0.6B/4B/8B + Stream 0.6B/4B/8B; 119 languages | Harm categories incl. jailbreak inputs | [arXiv 2510.14276](https://arxiv.org/abs/2510.14276) |
| **Granite Guardian** (IBM) | 2B / 3B-A800M / 5B / 8B, v3.0–4.1 + HAP classifiers | Input/output harm, jailbreak, RAG, hallucination | [arXiv 2412.07724](https://arxiv.org/abs/2412.07724) |
| **ShieldGemma** (Google) | Gemma 2 2B/9B/27B | 4 harm categories | [arXiv 2407.21772](https://arxiv.org/abs/2407.21772) |
| **WildGuard** (AllenAI) | Mistral 7B v0.3 | Prompt/response harm + refusal; taxonomy includes cyberattack | [arXiv 2406.18495](https://arxiv.org/abs/2406.18495) |

The broader guard field extends well past these five — OpenAI's policy-conditioned [gpt-oss-safeguard](https://openai.com) MoE (20B/120B), NVIDIA's [Aegis / NeMoGuard](https://arxiv.org/abs/2404.05993) line, and multilingual edge guards like WalledGuard, [HaloGuard](https://arxiv.org/abs/2607.02079), [PolyGuard](https://arxiv.org/abs/2504.04377) and [DuoGuard](https://arxiv.org/abs/2502.05163) — but the five above are the load-bearing ones for a general deployment. The size range is the deployment lever: a 0.6B Qwen3Guard or a 2B ShieldGemma is cheap enough to run on every turn, while an 8B Llama Guard or Granite Guardian is a heavier check you might reserve for the output side.

## The judges are not filters

One category in the taxonomy looks like a guard and is not. **[HarmBench classifiers](https://arxiv.org/abs/2402.04249)** (CAIS — Llama2 13B + Mistral 7B) are evaluation judges: they score whether an attack succeeded across a benchmark. They are explicitly *not* runtime filters. The same is true of the judge family more broadly — [ShieldLM](https://arxiv.org/abs/2402.16444) (Tsinghua) and [MD-Judge](https://arxiv.org/abs/2402.05044) (OpenSafetyLab, on SALAD-Bench) are built to grade transcripts, not to gate live traffic.

This matters in both directions. Putting a judge inline gets you a slow, offline-shaped model doing a job it was not calibrated for. And — the more common error — reading a guard's benchmark score as if it were a judge's leaderboard number conflates "how well it filters" with "how well it grades." When a security model's card reports a strong HarmBench figure only when *paired with an external Llama Guard*, as some do, that is the tell: the base model is not meant to be the sole safety layer, and the judge score and the filter are two different measurements.

## How the pieces stack

The honest deployment pattern is layered, because each of the three jobs above covers a gap the others leave open. For a serious deployment, the strongest arrangement is:

1. **A dedicated injection detector** on the input side — Prompt Guard 2, PIGuard, or Sentinel — to catch instruction-hijacking before it reaches the model.
2. **A separate content-policy guard** — Qwen3Guard, Llama Guard, or Granite Guardian — to check prompts and responses against a harm taxonomy.
3. **Task-specific cyber models kept behind operational controls** — tool permissions, sandboxing, and audit logs — so that a model's output cannot act on the world unmediated.

No single model in this survey does all three, and none of them is a substitute for the controls in step three. A detector reduces the rate at which an attack reaches your model; a guard reduces the rate at which harmful content leaves it; permissions, sandboxing and logging bound what happens when both are wrong. That layering — not any one classifier's score — is the defensive posture worth deploying.

This is the half of the ecosystem that the offensive catalog does not show. AdversariaLLM is being built to measure both, and the defensive layer is measured the same way as the offensive one: every model named traces to a card or a paper, the injection detector is never confused with the content guard, and the evaluation judge is never mistaken for a filter.
