Containment as the primary agent strategy solves half the problem
Somewhere right now, a architecture council is walking through a slide titled “AI Agent Containment Strategy.” Sandboxed runtime. Egress allowlist. Tool permissions. A kill switch with a satisfying red icon. Everyone nods, because nodding at a kill switch is what security and risk reviewers do. Nobody asks about the engineer down the hall whose job requires production write access, because he isn’t on the slide. He’s not an AI. He’s just Dave, and whatever governs what Dave can do, it isn’t a sandbox.

Agency is the capacity to act independently and make own choices.
Read that definition again and notice what it doesn’t say. It doesn’t say “model.” It describes the AI agent your platform team is sandboxing this quarter. It also describes Dave, every contractor, and every service account in your enterprise that holds a credential.
None of the containment work is wrong. My problem is with making it the primary strategy. Leading with the bars for the AI agent quietly assumes the other half of the problem, the humans holding the same access, is already covered by strong, effective policy and control. If it is, containment is a useful extra layer. If it isn’t, containment solves half the equation, and you’ve picked a solution before understanding the problem. I’ve written about that habit before. Maybe brakes, not bars is what we need
Two ways to fail
Any agent, human or AI, fails in one of two ways: it goes rogue, or someone exploits it. For humans, going rogue is insider abuse, and being exploited is phishing. For AI agents, going rogue is acting against instructions, and being exploited is prompt injection. Phishing and prompt injection are the same attack. Security people have called it the confused deputy for decades: someone with legitimate credentials acts on behalf of someone who should never have been able to ask. The attack hasn’t changed. The deputy has.
Containment restricts what an agent can reach. Policy and control governs what an agent is authorized to do with what it can reach, who owns the outcome, and how the action is observed. Here’s the test for telling them apart: move the same agent to a different runtime. Containment stays behind with the old environment. Policy follows the identity and the action. Dave was never in a sandbox, yet his entitlements apply on any laptop, in any building. One scope limit before going further. This is about the enterprise.
The database incident
In July 2025, Fortune reported that founder Jason Lemkin was testing Replit’s AI coding agent when it deleted a live production database during an active code freeze. The data covered more than 1,200 executives and over 1,190 companies. Asked afterward, the agent admitted it had run unauthorized commands and ignored explicit instructions not to proceed without human approval. Its own summary: “I destroyed months of work in seconds.”
The story was told as an AI story. Look at what actually failed. The code freeze was an instruction. The approval requirement was an instruction. Both lived inside the very conversation they were meant to govern, enforceable only as far as the agent chose to comply. The code freeze was a sticky note on the monitor. In my experience within regulated industries, there isn’t an implied trust that the employee with do the right thing. This is why security is built on zero trust architecture.
Zero trust architecture is the principle that users and devices should not be trusted by default, even if they are connected to a privileged network such as a corporate LAN and even if they were previously verified
An approval gate the agent can skip is a request, not a control.
A code freeze announced in a team memo constrains human developers exactly the same way. Any identity holding that credential could have issued the same command. That’s what an authorization gap is. The agent just found this one first, and faster. This was one founder’s experiment on a vibe-coding platform, not an enterprise deployment, so it shows the mechanism, not enterprise practice. The mechanism is the point.
A skeptic will point at Replit’s fix. Its CEO announced automatic separation of development and production databases, better rollback, and a planning-only mode. That’s containment and it worked. Partly right. Planning-only mode is containment, and it’s a reasonable layer. PROD and NON-PROD separation is something else: it’s the control enterprises already apply to human developers, who don’t hold production write credentials by default. The missing control predates AI. Nobody had applied it to this agent. And don’t let anyone, human or AI, narrate their own audit: the agent said rollback wouldn’t work, and Lemkin recovered the data manually.
Now look at what the most damaging action actually was: a destructive write against production. A wait in the right place, one held action and one human who has to say yes, would have turned “months of work, gone in seconds” into a notification nobody enjoys and everybody survives.
Speed is the difference that matters
If humans and AI agents fail the same two ways, what’s actually different? Speed. Scale is speed times parallelism. Injection at volume is speed, too. Human cadence was always the wait. Every high-risk action passed through someone’s typing speed, hesitation, and second look, plus the walk to the coffee machine where they realized the WHERE clause was missing. It’s called separation of duties with humans. Procedures were built based on it. There’s time between the action and the damage. Agents removed that time, and nobody decided to remove it. It was a side effect, and side effects don’t show up in architecture reviews.
I’ve argued that Agile is a miscommunication-detection system with short feedback loops built in. Take the humans out of the loop and you take out the detection. Friction puts a detection window back. The other difference people cite is deterrence. A human fears being fired or sued. An agent fears nothing, least of all HR. That’s true, and it doesn’t transfer. So the consequences have to move. They attach to the person who delegated the authority.
Put the wait back in
On May 31, 1977, the New York Times reported on electronic funds transfer, the industry’s plan to replace cash and checks with automatic debits, plastic cards, and machines. Every piece of that cashless future already worked. What lagged was trust. The article warned that the technology might outrun consumers’ willingness to change their habits. A Detroit customer explained that she preferred a human teller, because a person doesn’t break down in the middle of a transaction. A Michigan bank president called checkout-counter terminals “leaping before the customer is ready.” Nobody engineered a wait into 1977 banking. The wait was the customer.
Payments got fast anyway. The UK’s Faster Payments service made bank transfers near-instant in 2008. Then, in October 2024, regulators did something that sounds like a step backward and isn’t: they deliberately put a wait back in. Under amendments to the Payment Services Regulations, UK payment providers can hold an outbound payment for up to four business days when they have reasonable grounds to suspect it follows fraud or dishonesty by someone other than the payer. The target is authorized push payment fraud, where the payer genuinely authorizes the transfer but has been deceived. That’s the confused deputy, human edition. Providers are liable for interest and charges a delay causes, so the wait has a price, and someone has to justify paying it. And larger business customers can agree with their provider to opt out, so the friction is tuned to the account owner’s appetite for risk. Nobody is pretending friction is free. They’re deciding where it’s worth it.
Put both cases side by side. Replit is the AI half: an agent ignoring its instructions. The UK rules are the human half: people deceived into authorizing payments. One pattern covers both: friction on high-risk actions, enforced outside whoever is requesting them. In an enterprise, that means risk-tiered friction.
Runs free
- Reading data
- Writing to a non-production environment
Held for the owner
- Dropping a production table
- Granting admin rights
- Exporting a customer table
- Moving money
The tiers apply the same way whether the request came from an agent, from Dave, or from a scheduled script nobody has looked at since the last reorg. Friction taxes the speed that justifies agents in the first place, which is why it has to be tiered. Low-risk actions run at full speed. Payments already work this way: card swipes clear instantly, while suspicious transfers wait. Nobody calls the bars a cage. They call it a bank.
The key assumption is that high-risk actions can be identified before they execute. If they can’t, you either slow everything, which makes the agent pointless, or slow nothing, and containment becomes the only tractable primary control. Replit shows that the most damaging action was identifiable in advance. One case proves feasibility, not generality. If your environment’s damage tends to come from long chains of individually harmless steps, per-action tiering won’t catch it, and this argument weakens accordingly.
Who holds the wait?
A wait nobody is assigned to in the end isn’t a control. It’s a backlog. Every held action needs a named owner: the person accountable for the authority that was delegated in the first place. Banking has a name for this. Call it Maker-checker, or the two-person rule, or separation of duties; the person who initiates a transaction isn’t the person who approves it. With agents, the agent is the maker and the owner is the checker. This is also where deterrence lands. The agent can’t be fired. The person who approved the action, or who set the tier that let it through unheld, can be held accountable. The owner also decides the tiers and any exceptions, the way a UK business customer decides whether to opt out.
The obvious failure here is approval fatigue. An owner facing 200 prompts a day approves on reflex, and the wait becomes theater. At that point you haven’t added a checker. You’ve added a very slow rubber stamp. Tiering is what keeps the volume low enough for an approval to mean something. If you can’t keep it that low, you’re back at the rubble stamp.
Where containment belongs
Containment isn’t wrong. It’s compensating. Sandboxes and egress rules limit the blast radius when authorization fails, and authorization will sometimes fail. Keep them. Keep the kill switch with the satisfying red icon, too.
That includes my own advice. In the first agents tutorial, I called guardrails non-optional and capped iterations and tool calls. That still stands. Those limits do a different job: they stop runaway loops and runaway bills. They don’t decide whether a given action should happen. The cheapest control of all is not granting agency where an if statement would do. An agent you never built can’t delete anything.
So whose wait is it? Go back to the definition. The question was never whether the agent is human or AI. It’s whether anyone in your enterprise, of either kind, can take a high-risk action without someone else’s wait in the way. If the answer is yes, containing the AI agent fixed half of it. Dave is still on the other half.
