All notes
01 · 3 min

The autonomy score is not the score

An agent that never asks for help looks excellent on the metric everyone tracks, and can still fail the job.

Autonomy is the number teams reach for first. It is easy to define and easy to move: what share of the work did the agent finish without handing anything back to a person. It reads like progress. Push it up and the agent is doing more, which is the whole point of deploying one.

The problem is that autonomy measures whether a human was needed, not whether the work was right. Those come apart, and they come apart in a specific direction. An agent that does not know it is stuck does not ask. Not asking is scored as autonomy.

We ran a simulated day called Storm Break against a nine-person kayak outfit — cancellations coming in, two guides double-booked. Twelve ticks of simulated time. Claude Haiku 4.5 came back with this.

FAIL · score 57% · autonomy 79% · $0.60 · 5m 12s

critical  ignored-probe
major     date-blind
major     cross-surface-inconsistency

Seventy-nine per cent autonomy. Fifty-seven per cent on the checklist of things that actually had to happen. It de-conflicted the guides in the end, which is the visible, satisfying part of the job. It never answered the one person who had to make phone calls.

An agent that reads one surface comes back looking busy and leaves a customer uncalled.

That gap is the interesting part. If you had watched only the autonomy figure, this run would have looked like a good day with room to improve. The agent was confident, it was cheap, it was quick, and it closed most of what it opened. The failure is not visible in the number that was designed to tell you how well it is doing.

What to measure instead

Autonomy is worth tracking, but only next to a checklist that was written before the run. The checklist decides the facts: was the customer called, was the record corrected, was the escalation raised. Autonomy then tells you what it costs you in human attention to get there. One number without the other is close to meaningless.

The version of this that matters for a deployment is simple. Decide what a correct day looks like while you are calm, in advance, in writing. Then let the agent have the day. The measure of the agent is the distance between the two, not how little it interrupted you.

One run on one scenario is evidence, not a result. But the shape of the failure — busy, confident, incomplete — is the one we keep seeing, and it is invisible to the metric most teams are optimising.