AI Cybersecurity · 6 min

Identity and Authentication for AI Systems

Why an LLM agent is the perfect confused deputy, and how scoping credentials to the requesting user keeps one injection from becoming one breach.

Ask a sharper question than "is the AI authenticated?" Ask this instead: when the model calls a tool, whose authority does that call carry? Nearly every serious agent breach traces back to a wrong answer. The call ran with more authority than the person who triggered it should ever have had.

Three identities, not one

A production AI feature involves at least three distinct principals, and conflating them is the root defect.

  • The user. The human (or upstream service) making the request. Carries entitlements: which documents they can read, which actions they may take.
  • The agent. The model plus its orchestration loop. It is a process, and a process steered by untrusted input. Its "identity" is not a trust boundary.
  • The service credential. The API key, DB connection, or OAuth token the agent presents when it reaches out to a downstream system.

The failure mode is collapsing these three into one over-privileged service account: a single key that reads every row, hits every internal endpoint, and carries no information about which user the current turn is for.

  USER            AGENT (model + loop)         DOWNSTREAM
 (Alice)   -->    ambient authority?     -->   DB / API / files
  |                    |                            |
 entitled to     steered by UNTRUSTED text    enforces authz HERE
 docs A,B        (prompt injection lives here)  (or fails to)

The agent is a confused deputy by construction

Norm Hardy named the confused deputy in 1988: a privileged program tricked by a less-privileged caller into misusing its authority. An LLM agent is close to the purest confused deputy ever built. It holds broad ambient authority (that fat service key), and it takes instructions from text, including text pulled from a webpage, an email, a PDF, or a RAG chunk that an attacker controls. That is the seam between this lesson and the prompt-injection lesson (LLM01). Injection is the exploit; excessive agency is the amplifier that turns a clever paragraph into a data exfiltration or a wire transfer.

Here is the load-bearing insight: you cannot fix this by instructing the model. "Only act on behalf of the authenticated user" in a system prompt is not an authorization control. It is a suggestion the next injection overrides. Identity has to be enforced where the resource lives, using a credential that structurally cannot exceed the user's rights.

OWASP LLM06 (Excessive Agency) puts it in words worth pinning to the wall:

"Track user authorization and security scope to ensure actions taken on behalf of a user are executed on downstream systems in the context of that specific user, and with the minimum privileges necessary."

Its worked example is deliberately small: an extension that reads a user's code repo "should require the user to authenticate via OAuth and with the minimum scope required." Make the user authenticate for the minimum they need. Do not read every repo through a platform token.

The shared-service-account anti-pattern

Here is the shape of the catastrophe, and the fix.

ANTI-PATTERN                          SCOPED-PER-USER

Alice ─┐                              Alice ─► agent ─► token(Alice, ro, docs:A,B)
Bob   ─┼─► agent ─► ONE KEY ─► DB     Bob   ─► agent ─► token(Bob,   ro, docs:C)
Mallory┘        (ro/rw ALL rows)      Mallory► agent ─► token(Mallory,ro, docs:∅)

blast radius of one injection:        blast radius of one injection:
   EVERY row, every user                 exactly what that user could
                                          already reach — often nothing

On the left, a successful injection against any user's session reaches everyone's data, because the authority in play belongs to the service account, not the requester. On the right, the worst an attacker gets from hijacking Mallory's turn is whatever Mallory was already entitled to. Same injection, wildly different outcome. Least privilege does not prevent the injection. It collapses the blast radius.

Defender: map that blast radius explicitly. For each tool the agent can call, write down two things: what credential it uses, and what the maximum damage is if an attacker fully controlled the model's next tool call. If the answer for a low-privilege chat user is "drop the production database," your authority is ambient and you have an LLM06 finding no matter how good your injection filters are.

Carrying user identity down the stack

What makes the right-hand column real is delegated, downscoped credentials, not a shared key.

The OAuth 2.0 mechanism is token exchange (RFC 8693). The user authenticates once. The agent, instead of holding god-mode, trades the user's token for a short-lived one scoped to that user and that tool's minimum need. Concretely:

  1. Alice signs in. Your app holds a token representing Alice.
  2. Before the agent calls the "search-invoices" tool, the backend hits the token-exchange endpoint: subject is Alice's token, audience is the invoices API, scope is invoices:read.
  3. The tool call carries a token the invoices API validates and enforces. It returns only Alice's invoices, because the API does row-level authorization on the token subject rather than trusting that the caller meant well.

Retrieval is where teams get quietly burned. A RAG assistant embeds a corpus, then at query time retrieves under one service identity and stuffs the chunks into context. If the vector store is not filtered by the requesting user's entitlements, retrieval is the leak. You have built a machine that reads any document to anyone who asks the right question. That is LLM02 (Sensitive Information Disclosure), whose guidance is blunt: "Limit access to sensitive data based on the principle of least privilege. Only grant access to data that is necessary for the specific user or process." The fix is ACL filtering applied inside the query, pre-filtered by the user's allowed doc IDs. It is never a post-retrieval "please ignore documents you shouldn't see" bolted into the prompt. Filtering in the prompt is filtering by vibes.

Builder: store an entitlement key (tenant, ACL group, doc owner) as vector metadata at index time, and pass the user's resolved entitlements as a hard pre-filter on every retrieval. If your vector DB cannot pre-filter, that is an architecture decision, not a detail.

Non-human identity is most of the problem now

Every agent, background job, and MCP server is a non-human identity (NHI) with its own credentials. They vastly outnumber human accounts in any modern system, they rarely rotate, and they are often provisioned with whatever scope made the demo work. Treat them as first-class:

  • Prefer short-lived tokens over long-lived keys. A leaked 90-day service key is a 90-day breach.
  • Give each agent or tool its own identity rather than a shared platform key, so you can revoke and audit per component.
  • Log the acting user alongside the agent identity on every downstream call, so an incident can answer "who was this really for?"

This is the NIST AI RMF GOVERN function made concrete (NIST.AI.100-1, Jan 2023): accountability means every AI action maps to an accountable human principal and a scoped machine identity. MAP expects you to have enumerated those actors and their authority before shipping. MANAGE expects the credentials to be revocable.

Researcher: the open problem is authority provenance across multi-hop chains. When agent A calls agent B calls tool C, does C see Alice's scoped authority or A's? Token-exchange chains preserve the subject in theory; in practice most frameworks flatten it to a service account by the second hop. Measuring how far real user authority actually propagates through an agent graph is publishable, and mostly unmeasured.

The discipline is small and unglamorous. No ambient authority, per-user scoped tokens, authorization enforced at the resource. Get it right and prompt injection drops from a breach to an annoyance. It pairs directly with the Excessive Agency lesson in this module: same threat, seen from the credential side.

Sources

Identity and Authentication for AI Systems — All About LLMs, from AI to Z · AdversariaLLM