Done, Keys, and a Second Check
TLDR
Vendors sell task-completion. Enterprises buy without owning outcomes, keys, or mistakes. One sandbox is not a control system (Noam Brown). The missing work is roles, workflows, decision rights, and controls: who defines done, who holds the keys, who may escalate or monitor, and who owns the second check.
The root cause: treating probabilistic agents like deterministic SaaS
Enterprises treat probabilistic agents like deterministic SaaS. They copy the consumer “build something agents want” form factor (Levie, 19 Sep 2026) and skip ownership. That shape becomes a shopping list: APIs, MCP servers, connectors, sandboxes, spend wallets. Procurement buys every item. The handoff has no owner.
Who decides the task is complete? Who authorised the keys? Who stops the run when severity flips? Those answers do not ship with the SDK.
Build something agents can use. Then design something humans can own.
When the blast radius has a name on it
Pick GDPR deletion. The agent triages customer requests. “Delete my account” comes in. The agent classifies: routine, low-severity, auto-execute.
The customer meant “pause my subscription.” The agent read “delete.” It has write access to the CRM. It deletes the account. The customer is a £2M enterprise client mid-renewal.
Escalate queue: The agent routed edge cases to “escalate.” No named owner. No SLA. The queue has 400 items. This one sat for six days.
Keys without a grantor: The agent has CRM write because “end-to-end completion” shipped without least privilege. Who approved that scope? Who can revoke in under an hour?
Who gets fired: Not the agent. The compliance officer who signed off? The product owner who built it? The vendor who sold task-completion?
Keys without a named grantor and a revoke path are blast radius with a logo.
Done, keys, and triage are the real product
Levie’s Jev × Box demo shows judgement at machine speed. Pull an incident. Ask if it is customer-facing and how severe. Route to escalate, monitor, or review. Stamp metadata. Fast. Cheap (18 Sep).
The folders are the easy part. The rights are not.
Done. Escalate, monitor, and review are verdicts, not labels. Someone must say when monitor is allowed and when escalate is mandatory. Someone must define done for the person who inherits the file. Otherwise the agent classifies, and the incident still has no owner.
Keys. End-to-end completion means write paths: systems of record, spend, customer data. Consumer agents want transactions. Enterprise agents want access. Muse-style handoff needs tools and the ability to transact (Levie on Muse). That works when the agent competes for attention, not when it deletes the renewal pipeline.
Write access is not symmetric. Draft-write is reversible: a human commits. State-write has immediate structural impact: CRM state changes, schema migrations, data shares, payments execute. True autonomy leans on state-write. That is why the governance argument is mandatory.
Triage rights. The taxonomy only works if escalate, monitor, and review have named roles, time boxes, and an audit trail. Else the agent sorts into queues nobody empties.
Heroes still need this. Connectors and consolidation are not enough (From 100 Agents to Two Heroes). Without verdict rights, a hero is a pilot with better branding.
One fence is not a control system
Giving a day-one intern a corporate card with a £5,000 limit does not eliminate fraud risk. It caps the ruin. The control is not the limit. The control is counter-signature, monthly reconciliation, and a manager who can freeze the card.
A sandbox works the same way. It caps waste. It does not prevent it. One fence is single-point failure dressed as defence.
Noam Brown’s clarification matters. Absolute isolation guarantees are hard. Layers of defence matter. Agents that are supposed to be independent can still coordinate with very few bits. The Hugging Face lesson, in his words: too much trust in the sandbox, not enough independent safeguards. Design as if you overestimated the risk (Noam, 18 Sep; Dwarkesh episode post).
Dwarkesh puts the operator question next to the lab one. When capability rises, how do you know the system is aligned before you grant more autonomy (18 Sep)?
Enterprises hear “sandbox” and relax. That is single-fence overconfidence. A sandbox is one control. A control system stacks policy, least privilege, independent monitoring, human stop authority, and a second check the deploying team does not own.
Anca Dragan and Rohin Shah argue we should keep Chain-of-Thought monitorable on purpose, not assume it lasts (Anca on X; DeepMind Institute essay). If you cannot see how the agent reasoned, your “monitor” path is a folder name.
Stop authority: the circuit breaker and who owns it
Pick one pattern. Make it concrete. Visualise the circuit breaker.
Pattern: A human-in-the-loop dashboard with a deterministic financial threshold. Any action above £10,000 cumulative impact routes to mandatory human review before execution. The dashboard shows agent reasoning, proposed action, and calculated risk score.
Owner role: The Compliance Lead, independent of the deployment team, has stop authority. They can halt the workflow, revoke agent keys without escalation, and approve or reject state-write actions before execution. Audit trail required. Decision logged and reviewed quarterly.
For critical systems, pre-execution gates matter more than post-execution rollback. You cannot un-ring the bell. A CRM delete, an unapproved vendor payment, a schema migration: the damage happens on commit. A four-hour reverse window is an illusion for state-write. Build the circuit breaker before the write, not after.
That is one version. The pattern matters less than naming it, funding it, and giving one person outside the product team the power to say no.
Four answers before go-live
Pick one workflow. Answer these in writing before you ship.
- Done: business definition for that decision class, not model scores.
- Keys: what the agent may read or write, who grants, who revokes, what stays human.
- Triage: who can place work in escalate, monitor, or review; who clears those queues; under what SLA.
- Second check and stop: which independent role can halt or reverse, with an audit trail.
Vague answers mean you bought a form factor. You did not ship an operating model.
Sources
- Aaron Levie, Personal agents and “build something that agents want” (Muse form factor), 19 Sep 2026.
- Aaron Levie, Jev × Box demo: escalate / monitor / review routing, 18 Sep 2026.
- Noam Brown, Layered defence and over-trust in sandbox isolation, 18 Sep 2026.
- Dwarkesh Patel, Episode with Noam Brown (multi-agent, HF, alignment before RSI), 17 Sep 2026.
- Dwarkesh Patel, How will we know models are aligned before RSI?, 18 Sep 2026.
- Anca Dragan, Preserve Chain-of-Thought monitorability, 16 Sep 2026; Rohin Shah and Anca Dragan, The case for reasoning transparency, DeepMind Institute.
Challenge
Pick one agent workflow you call production this week.
- Who, by name, defines done?
- If the agent goes rogue, who holds the kill switch, and how many minutes does it take them to flip it?
- Who owns the second check, and who has stop authority outside the deploying team?
If you need a slide to answer, you copied what agents want. You did not ship an operating model.