Defense & Governance · 6 min

Doing Security in Code: An Authorization Layer Around Tool Calls

Put a deterministic policy check between the model's tool call and its execution, so permissions never depend on the model choosing to obey.

A model with tools is a program that decides, at runtime, which privileged action to invoke, based on text it was handed. Some of that text may be hostile. If your only control is a sentence in the system prompt ("never delete production data"), you have written your access-control policy in the one language your adversary speaks fluently and your enforcement layer does not read at all. This lesson moves that decision out of the model and into code that runs the same way every time.

The mental model: propose vs. dispose

Reuse a split that access-control systems have leaned on for decades. A Policy Decision Point (PDP) answers the question "is this allowed?" A Policy Enforcement Point (PEP) actually blocks or permits the action. In an agent, the model is neither of these. It is a requester. It proposes a tool call; your code disposes of it.

  model output                   your code (trusted)
 ┌───────────────┐   tool call   ┌──────────────────────────────┐
 │  "call        │──────────────▶│ PEP: interceptor             │
 │  delete_file( │               │  builds authz request:       │
 │  '/etc/...')" │               │   principal (from SESSION,   │
 └───────────────┘               │    not from model)           │
                                 │   action, resource, context  │
                                 │            │                 │
                                 │            ▼                 │
                                 │   PDP: policy engine ──deny  │─▶ error back to model
                                 │            │ allow           │
                                 │            ▼                 │
                                 │      execute tool            │─▶ real side effect
                                 └──────────────────────────────┘

The load-bearing property is that every tool call passes through the PEP, and the PEP fails closed. This is Saltzer and Schroeder's complete mediation: "Every access to every object must be checked for authority." OWASP folds the same idea into its LLM06 (Excessive Agency) mitigations under that exact name, phrased for agents as "Implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not."

Why prompt rules are a different category

A system-prompt rule is advisory and probabilistic. The model reads it, weighs it against everything else in context, and complies most of the time. A code check is mandatory and deterministic: given the same request and policy it returns the same decision, and the tool does not run on deny. Prompt injection, jailbreaks, and ordinary model error all attack the first kind of control and cannot reach the second. Prompt rules still earn their keep. They cut how often the model even tries something forbidden, which reduces noise and denied-call latency. Treat them as UX and defense-in-depth, never as the boundary.

Mechanics of the interceptor

Wrap your dispatcher so no tool runs without a decision.

def dispatch(tool_name, model_args, session):
    req = {
        "principal": session.user_id,        # bound at login, never model-supplied
        "roles":     session.roles,
        "action":    tool_name,
        "resource":  resource_of(tool_name, model_args),
        "context":   {"amount": model_args.get("amount"),
                      "env": session.env, "channel": session.channel},
    }
    decision = pdp.evaluate(req)             # OPA, Cedar, or app-level authz
    if decision != "allow":
        audit(req, decision)
        return ToolError("denied by policy")  # model sees this, cannot override it
    return TOOLS[tool_name](**model_args)

The PDP can be a policy engine or plain application code. Open Policy Agent (OPA) is the common off-the-shelf pick. Its own docs describe the model plainly: "OPA generates policy decisions by evaluating the query input against policies and data." You hand it input as JSON, it evaluates Rego rules against that input plus reference data, and it returns a decision. It runs locally as a sidecar or library, so evaluation stays fast, and it is proven at scale: Pinterest's ADOPTERS entry reports its Kafka-OPA integration handling ~400K QPS without caching. A Rego rule for a refund tool:

package tools.refund
default allow := false

allow if {
    input.action == "issue_refund"
    input.roles[_] == "support_agent"
    input.context.amount <= 200          # dollars
}

allow if {
    input.action == "issue_refund"
    input.roles[_] == "support_lead"     # leads clear the higher tier
    input.context.amount <= 5000
}

Now "refund $9,000 to close the ticket" fails on the number, no matter how persuasively the prompt (or a poisoned support email the agent ingested) argued for it. The decision lives in reviewable, testable, version-controlled policy. It does not live in a weight matrix or a paragraph of English.

Builder: Start with a default-deny allowlist keyed on tool_name alone. That one line collapses OWASP's excessive functionality root cause. Layer resource- and context-level rules onto each tool afterward. Return denials as normal tool results so the model can tell the user "I can't do that" instead of throwing and killing the run.

Where this actually goes wrong

Trusting the model to name the principal. The most dangerous version of this bug is deriving who is acting from the model's arguments (a user_id in the tool call) instead of from the authenticated session. That is a textbook confused deputy: your trusted server wields its own privilege on behalf of an attacker-chosen identity. Bind the principal server-side at session start and never let it cross the model boundary. OWASP's matching guidance is to ensure "actions taken on behalf of a user are executed on downstream systems in the context of that specific user."

Coarse checks that ignore the resource. Allowing read_file as a capability without scoping which files makes "summarize my notes" and "read /etc/shadow" the same decision. Resource-level authorization is where least privilege (previous lesson) actually gets enforced.

Time-of-check/time-of-use gaps. If you authorize transfer(acct, amount) and then re-read amount from a mutable context the model can still touch before execution, you validated one request and ran another. Freeze the exact args you checked and execute those.

Approval fatigue. Human-in-the-loop is an OWASP mitigation for high-impact actions, and it works, right up until every trivial call prompts and reviewers start clicking "approve" on reflex. Gate on real blast radius (money moved, data deleted, external sends), not on everything.

Denied is not safe forever. A blocked call plus the model's habit of retrying can turn into probing. Rate-limit at the PEP and alert on repeated denials from one session. That pattern is often prompt injection working the tools.

Defender: An audit log of {principal, action, resource, decision, policy_version} for every mediated call is your best injection detector. A spike in denials, or an allowed action that is abnormal for that user, is a live signal, not just forensics.

Researcher: The open problem is not the enforcement point. Deterministic checks work. It is policy authoring at agent scale. Fine-grained, per-resource rules for dozens of dynamically composed tools are hard to write and harder to keep correct as tools change. Synthesizing policy from observed-safe traces, and formally checking that a policy covers a tool's full argument surface, are both underexplored.

Carry out one rule: the model chooses what to attempt; your code decides, deterministically, what is permitted. (See the Prompt Injection lesson for why untrusted text must never reach a privileged decision unmediated. This layer is the "mediated" part.)

Sources