Evidence unlocks permissions, not better models
The constraint on agent autonomy stopped being capability a while ago. It is warrant.
There is a version of the agent roadmap where everything waits on the next model. The agent cannot be trusted with refunds above a certain size, so you wait. The next release lands, it is better at reasoning, and you raise the limit a bit. Progress is something that happens to you when a lab ships.
That is not usually what is happening. In most cases the agent is already good enough to do more than it is allowed to do. What is missing is not capability. It is anything that would justify widening the permission.
Evidence, not a better model, is what unlocks the next permission.
This matters commercially, because the permissions are where the money is. An agent that drafts a reply for a human to send saves a little time. An agent that sends it removes a person from the loop. Each step up in autonomy is worth real revenue, and each step is gated by somebody being willing to sign off on it.
Who actually signs
The person blocking the step is often not technical, which changes what counts as evidence. A common pattern is that non-technical stakeholders do not understand what the agent will and will not do, so they will not approve launch. The buyer for assurance is not the machine learning engineer. It is the executive who has to carry it.
Those people are not moved by a better benchmark average. They are moved by being able to see the specific situations that worried them, and what the agent did in each. The question in their head is not whether the model is good. It is whether this agent, doing this job, will do something they will have to explain.
So the artefact that unlocks the permission looks less like a score and more like a record. Here are the situations we were afraid of. Here is what it did. Here is where it stopped and asked. That is legible to someone who has never read an eval report, and it is the thing a partner bank or an auditor is really asking for when they ask how the agent is tested.
Once you see it this way, testing stops being a cost centre. It is the thing standing between the agent you have and the agent you are allowed to run.