"Ask the User to Approve" Is Not a Security Control
Approval prompts get clicked through by the tenth repetition. Treat human review as a scarce resource and spend it where the blast radius is — everywhere else, enforce in code.
Most agent products ship with the same safety story: the agent can do risky things, but it asks first. A dialog appears, a human reads it, a human clicks Allow. Review boards sign off on this. Threat models list it as the mitigation.
Now watch what happens on day three. The prompt that was a careful decision on Monday is a reflex by Wednesday. Nobody got lazy. The design asked a person to make the same judgment fifty times a day, with almost no signal about which fifty-first one matters.
That is alert fatigue, and we already know how it ends. The difference is that a SOC analyst ignoring a noisy rule misses a detection. A developer clicking through an approval dialog is the control, and it just stopped working.
A confirmation prompt is a control only while the human reading it is still deciding. Past that point it is a logging event with a button on it.
Why approvals decay
Four things compound.
Base rates. Nearly every action the agent asks about is fine. A reviewer
who has seen forty harmless npm install prompts is trained, correctly, to
expect the forty-first to be harmless too. The rare malicious one arrives
looking exactly like the common benign one.
The prompt hides the thing that matters. An approval dialog shows a tool name and arguments. It rarely shows why the agent wants this, what content it just read that pushed it there, or what the call touches downstream. In an indirect prompt injection, the dangerous part is the context that steered the call, and that is precisely the part the dialog omits.
Approval is cheap to give and expensive to refuse. Allow keeps the agent moving. Deny stalls the task, forces a rewrite, and makes the human the bottleneck. Every incentive in the loop pushes toward yes. Products then add an "always allow" checkbox to relieve the pressure, which converts a per-action control into a standing grant the first time someone ticks it.
The attacker gets to shape the request. An adversary who controls what the agent reads can phrase the action to look routine, split it into small innocuous steps, or time it for the middle of a long task when attention is lowest. Social engineering against the reviewer is cheaper than any exploit.
What human review is good for
None of this means remove the human. It means stop treating the human as a generic filter and start treating their attention as a budget with a very small balance.
Humans are good at a narrow set of judgments: Is this the thing I intended? Does this scope make sense for this task? Should this ever happen at all? They are poor at scanning a stream of near-identical technical requests for the one that is subtly wrong. Spend the budget on the first kind.
Put the boundary in code, then ask humans about the remainder
The practical pattern is three tiers, and the order matters.
1. Deny by construction. Things the agent should never do get no prompt at all, because a prompt implies a choice that doesn't exist. Egress to hosts outside an allowlist, reads of credential files, writes outside the working tree. Containment belongs to the environment, and a boundary the runtime enforces can't be talked past or clicked through. Schema and policy checks on tool calls do the same job one layer up: malformed or out-of-policy calls fail before a person ever sees them.
2. Allow by policy. Low-impact, reversible, well-scoped actions run without a prompt, under a narrow delegated identity rather than your full credentials. Reading files in the repo, running the test suite in a sandbox. This is where most of the volume lives, and removing it from the human's queue is what keeps the queue meaningful.
3. Ask, with context, for what's left. The remainder is actions that are consequential, irreversible, or cross a trust boundary: sending data outside, deploying, deleting, spending money, granting access. These are rare enough to earn real attention if you design the prompt to deserve it:
- State the effect, not the command. "Send the contents of
customers.csvto an external address" is a decision.curl -X POST ...is a string. - Show the provenance: what the agent read before proposing this, and whether any of it came from untrusted content.
- Make the default deny, and make approval scoped to this one action. No session-wide "always allow" for the dangerous tier.
- Rate-limit your own prompts. If a workflow generates more than a handful of high-tier asks per session, that is a design bug to fix, not a human to blame.
Measure the control, don't assume it
If approval is a control, it has a failure rate, and you can estimate it.
- Approval rate per prompt type. A prompt approved 99% of the time is telling you it should be a policy, not a question. Promote it to tier two or delete it.
- Time to decision. Median decisions measured in a second or less on dialogs that require reading are rubber stamps.
- Canary prompts. Periodically inject a clearly-bad synthetic request in a test environment and see whether it gets approved. It's the same idea as a phishing simulation, applied to your own review loop.
- Standing grants. Count how many users have "always allow" enabled and on what. That number is your real policy, whatever the documentation says.
Everything the reviewer decides should also land in your audit trail, with the context the human saw. When an approved action turns out to be the incident, "the user clicked Allow" is where the investigation starts, not where it ends.
The test
Take your agent's approval flow and ask one question: if the person reviewing it were tired, rushed, and had approved the last thirty requests, would the control still hold? If the answer is no, you have a speed bump, not a boundary. Move the enforcement into the environment and the policy layer, and keep the human for the decisions only a human can make.
If you want a quick look at what your own agent setup allows without asking, the agent config checker is a reasonable place to start.
Read next — more in Defense engineering
Your Voice Is Not a Password Anymore
Three seconds of audio from an old video is enough to convincingly be someone you love. The check that works doesn't happen on the call.
7 min read
Threat Modeling for AI Agent Systems
Threat modeling fails because it runs as a quarterly workshop. Run it in sprints, at the architecture layer, and start with the agent surfaces scanners can't see: tool calls, memory, MCP servers.
17 min read
Threat-Informed Defense Is Detection Engineering, Minus the Guessing
A detection built on a file hash lasts until the next build. One built on a technique costs the adversary something to route around. That's the whole argument.
5 min read