Defense & Governance · 6 min

Human-in-the-Loop: Gating Irreversible and High-Stakes Actions

When an agent can move money, delete data, or send mail, the approval gate is a real security control, but only if it fights rubber-stamping instead of manufacturing it.

An LLM agent that can call tools has agency: it turns text into effects in the world. Usually that is the whole point. The failure mode OWASP catalogs as LLM06: Excessive Agency is what happens when those effects outrun anyone's ability to catch a bad one, because there are too many tools, scoped too broadly, firing too autonomously before a human sees them. Human-in-the-loop (HITL) approval is the specific mitigation for the "too autonomous" axis. Think of it as a circuit breaker on the tool boundary, and like any breaker it only helps if it actually trips and someone actually reads it.

The mental model: gate the effect, not the model

Put the gate at the tool call, not inside the model's reasoning. The LLM proposes; a separate, non-LLM control decides whether the proposal executes. This matters because the model is the untrusted component. Prompt injection, hallucinated parameters, and a confused-deputy handoff all live inside its reasoning. If the thing deciding "is this allowed?" is also the thing that can be talked into anything, you have no control at all. OWASP is blunt about it: implement authorization in downstream systems rather than relying on an LLM to decide whether an action is allowed.

  user ── prompt ──▶ [ LLM agent ]
                        │ proposes: refund(order=8891, amount=$4,200)
                        ▼
                  ┌───────────────────────┐
   policy engine ─▶  is this high-stakes? │──no──▶ execute
   (NOT the LLM)  └───────────┬───────────┘
                              │ yes
                              ▼
                     [ approval queue ]  ◀── human reads full context, approves/denies
                              │ approved (signed, bound to exact params)
                              ▼
                     downstream system re-checks authZ, then executes

The gate is a policy decision (which classes of action need a human) enforced by infrastructure (the queue, the signed approval, the downstream re-check). The model's opinion about whether something is risky is an input to that decision. It is never the verdict.

Which actions must not auto-execute

The useful axis is reversibility multiplied by blast radius, not "sensitivity" in the abstract. Ask two questions. If this fires wrongly, can I undo it? And how far does it reach?

                 LOW blast radius        HIGH blast radius
  IRREVERSIBLE   send one email          wire transfer, prod DB DROP,
                 (embarrassing)          delete-all, publish to 40k users
                 → gate on anomaly       → ALWAYS gate

  REVERSIBLE     update a draft,         bulk tag 10k records,
                 toggle a flag           reversible-but-tedious rollback
                 → auto-execute          → gate above a threshold

The actions that belong in the always-gate column are concrete: payments and refunds, outbound sends (email, SMS, social, anything customer-facing), deletions and destructive migrations, permission and credential changes, and anything that spends money or provisions billable resources. OWASP's own examples sit here too, calling out email forwarding, mail deletion, and document deletion as the operations that need a human before they run. Everything reversible and small should auto-execute. Gating it is how you train people to stop reading, which is the real problem.

The failure mode is rubber-stamping, not missing gates

Most teams add the gate. Then it degrades into a click-through, and now you have something worse than nothing: the audit log says "human approved" while no human engaged. Two well-studied mechanisms drive this.

Automation bias. Skitka, Mosier and Burdick (1999) split operator error into two kinds. Commission errors mean following an automated recommendation that turns out to be wrong. Omission errors mean missing an event because the automation never flagged it. Both rise when a human is asked to "monitor" an automated agent. A reviewer staring at a stream of approve recommendations from a confident model does not become a skeptic. They become a rubber stamp. NIST's AI Risk Management Framework treats this under its Human-AI Configuration category and is explicit that oversight roles have to be designed against over-reliance, not merely assigned to someone.

Fatigue. The SOC-alert world is an honest preview of where a noisy approval queue ends up. Orca's 2022 alert-fatigue survey found around 59% of respondents fielding more than 500 cloud alerts a day, and roughly 55% reporting that critical alerts slip through, often weekly or daily. Volume is the enemy of attention. An approval queue that pages someone forty times a shift is a queue whose fortieth item gets the same reflexive approve as a cookie banner.

So here is the design consequence that most teams miss. Every gate you add for a low-stakes action is a tax on the attention available for the high-stakes ones. Gating is a budget, not a default.

Designing an approval that resists the stamp

A good approval payload surfaces enough to make deny a live option. Compare:

BAD  →  "Agent wants to run delete_records. Approve? [Y/N]"

GOOD →  ACTION      delete_records  (IRREVERSIBLE · no soft-delete)
        SCOPE       1,204 rows in `customers` WHERE last_seen < 2024-01-01
        SAMPLE      #8891 acme.co · #8892 globex.io · #8893 …  (12 shown)
        TRIGGERED   user msg: "clean up stale accounts"   ← show the origin
        REVERSIBLE  NO. Nightly backup is 19h old.
        COST/BLAST  affects 3 downstream billing jobs
        [Deny]  [Approve — re-auth required]

A few rules make the difference:

  • Bind the approval to the exact parameters that will execute. Approve these 1,204 rows, not "a deletion." That closes a time-of-check/time-of-use gap where the agent re-plans between approval and execution. The signed approval should carry a hash of the concrete call.
  • Show provenance. Which user turn, or which retrieved document, caused this? If the trigger is text from an untrusted source, that is your injection tell.
  • State reversibility and cost honestly, computed by the system rather than narrated by the model.
  • Make deny cheap and approve slightly effortful for the top tier, with a re-auth or a typed confirmation on irreversible actions. Friction on the dangerous path and not the safe one is the entire game.
  • Batch and rate-limit so ten near-identical proposals arrive as one reviewable group. Rate-limiting doubles as an LLM06 damage-limiter: it caps how many bad actions a compromised agent gets off before anyone notices.

Builder: Default new tools to requires_approval: true and make removing the gate the reviewed change. That is cheaper than discovering an ungated send_email after it has mailed your list.

Defender: Alert on the approval rate per reviewer. A 99% approve rate with sub-second decision latency is a rubber stamp with a paper trail. Treat it as a failed control, not a healthy one.

Researcher: The open question is calibration. Can the gating threshold be learned from action metadata (reversibility, blast radius, injection signals) without the model gaming it? Anything the LLM can influence, the LLM can be injected to influence.

What still goes wrong

Even a clean gate leaks. Aggregation: each send_email looks fine, but 5,000 of them is a spam incident, so gate on rate, not only per call. Opaque parameters: an approver cannot validate a base64 blob or an anonymous resource_id, so render human-legible summaries or the gate is theater. The confused deputy: if approval runs with the agent's privileges rather than the requesting user's, a low-privilege user gets high-privilege effects rubber-stamped, so execute in the user's security context with minimum privileges, as OWASP requires. And remember the gate is not authorization. The downstream system must independently re-check policy at execution time, because an approval token is not a capability grant.

This connects straight back to Module 6's work on prompt injection and tool-call provenance. The approval gate is where an injection attempt becomes visible to a human, but only if the payload shows where the instruction came from. A gate that hides provenance just hands the attacker a laundered, human-signed action.

Sources

Human-in-the-Loop: Gating Irreversible and High-Stakes Actions — All About LLMs, from AI to Z · AdversariaLLM