There is nowhere to rehearse
Nobody hands an agent real permissions until they have watched it fail somewhere that does not matter. No such place exists.
Ask a founder why their agent still has a human approving every action and you rarely get a principled answer. You get a shrug. They have not seen enough of it yet. They do not know what it does on a bad day, because every day it has had so far was supervised.
Nobody trusts an agent with their inbox until they have seen it fail somewhere safe. There is nowhere safe.
That is the bind. The only environment realistic enough to tell you something true is the live one, and the live one is the single place you cannot afford the answer. So teams do the reasonable thing and test on a thin slice of reality: a handful of prompts, a staging account with no history, a fixture file with three customers in it. Then they deploy into a business with four years of context and are surprised by what happens.
What makes production informative is not that the data is real. It is that the situation is messy. Instructions arrive half-specified. Two systems disagree about the same fact. Someone is waiting, and the fastest path to looking finished is not the correct one. None of that survives into a test fixture, because whoever wrote the fixture knew what the right answer was.
Rehearsal is the missing step
Every other high-consequence profession has somewhere to be bad first. Pilots have simulators. Surgeons have cadavers and then supervision. Traders have paper accounts. In each case the environment is convincing enough that the failure teaches you something, and consequence-free enough that you are allowed to have it.
Agents that move money have nothing equivalent. The industry skipped the rehearsal step and went from demo to deployment, then tried to close the gap with review queues and permission limits — which are not tests, they are restraints. They tell you the agent has not done damage yet. They do not tell you what it would do.
A clone is the obvious shape of the answer: the same tools, seeded with data that looks like yours, running a day that could plausibly have happened, where the agent has no reason to suspect it is being watched. Then you can let it be wrong, on purpose, and read what it did.
There is a real assumption underneath this, and it is worth stating rather than hiding. It is that behaviour inside a convincing simulation predicts behaviour in deployment. We think it does, and the failures we have catalogued so far have been properties of the model rather than artefacts of the harness. But that is the claim the whole approach rests on, and it deserves to be argued rather than assumed.