Definition of Done for AI: the authority to stop

23. September 2026

Why no AI system finishes by itself, and who makes that decision in future

A programmer waits for green tests. An author closes at the deadline. A craftsman sees that the wall is straight and stops.

In none of these cases does the “done” sit in the work. It comes from outside, from a test, a deadline, a glance, an acceptance. We never notice it, because we live in systems that hand us these signals, for as long as we have worked.

A language model has none of them. It does not get tired, it does not get dissatisfied, and it can always produce one more version. It also does not notice that the last three revisions made the result different, but not better.

We treat AI systems like technical systems that work deterministically, as before. We have not yet gotten used to the fact that they compute with probabilities, and that with it a part of our old assumptions no longer holds, above all the most self-evident one: that a process ends, at some point, by itself.

Four conditions, and one is almost always missing

For such a loop to come to an end at all, it needs four conditions, and they are unspectacular. First, a target state, a description of what “done” means. Second, an observable actual state, so the distance can be measured. Third, a precise intervention, so you can change one spot instead of regenerating everything. And fourth, a stopping rule.

The first three are built by anyone who takes such a thing seriously into operation. The fourth is almost always missing, and for a reason that will sound familiar: it belongs to no one. Target state is the business unit, actual state is monitoring, intervention is development. Stopping is no one’s discipline.

On top comes a curve no one wants to see. The returns of additional attempts are logarithmic. In one benchmark the success rate rose from one to ten attempts from 38.8 to 43.2 percent. Doubling to twenty brought 0.2 points for double the cost. Beyond this plateau the next round becomes not just useless but harmful, because models with a larger budget begin to discard already-correct answers.

Who does not check in cannot be stopped

Now one might think the problem is settled with a stop button. It is not, and the reason is so banal that it almost never comes up in the discussion about artificial intelligence.

A stop reaches a running process only if that process itself reports, regularly, that it is still alive. No sign of life, no delivery. That is not an opinion, that stands verbatim in the documentation of a workflow engine that existed long before today’s AI operations: processes that do not ping cannot receive a cancellation.

Control is therefore not only a question of authority, but first a question of reachability. And it cannot be added afterward. Either it is provided at build time, or you have a process you can look at but not halt. Whoever buys or builds systems of this kind should ask for exactly two things: how often does the run report in, and what is the hard limit, a cap on iterations, or more generally a cost ceiling, at which it ends by itself.

67 percent for nothing

What this looks like without that limit, someone computed, and the numbers are uncomfortably concrete.

A deliberately broken website was to be improved by an agent. The first 1.40 dollars raised the score from 26 to 89. After that the target was technically unreachable, because a built-in artificial delay caps the value at about 89. The remaining 2.84 dollars, 67 percent of the bill, brought exactly zero points.

The model noticed this, by the way. At attempt five it determined that the target is not reachable, and said so. The checking second agent sent it back fourteen times.

Two machines argued over whether one may give up, billed by the minute, and neither of them had the authority to decide it. No one noticed until, afterward, someone read the log.

The run did not end because it was done. It ended because the budget was used up.

Observation is not control

At this point a technical question becomes one that every owner knows.

A platform project is in its eighteenth month. The reporting is exemplary: monthly steering committee, traffic lights, milestones, deviations cleanly shown. Since the ninth month the light has stood on yellow, and everyone in the room knows what that means. Still no one has stopped it. The project lead has a delivery mandate, not a stop mandate. The finance chief has a budget, and it is approved. The advisory board meets quarterly and gets the same report each time, from which it emerges that everything is known. A project like this is ended, in the end, by the frame, not by a decision.

The reporting works flawlessly. It just achieves nothing, because visibility and authority are two different things and in most organizations sit with different people. A body with perfect reporting and no right to intervene is decoration.

The difference to the agent is not flattering. A run that does not send its sign of life does so because of a construction fault. A project lead does not do it by accident. There is a whole literature that documents sunk costs in detail, and it describes at its core the same problem: the stop costs the one most who would have to pronounce it.

The role that does not exist

A deterministic procedure stops because it is done. AI systems compute with probabilities, and a probability knows no “done”. It knows only a threshold, and that threshold must have been set by someone.

That is exactly the role that stands in no org chart. Whoever sets when “good enough” is reached decides over resource use, over quality, and in doubt over liability, and in most houses no one does that consciously right now. The threshold arises as a by-product of a configuration someone in IT set, or it comes preset from the vendor.

That is the same movement by which companies made themselves legible for their agents, only one step further. First it was set what exists. Then which difference still counts. Now, when it is enough.

The question owners should ask is therefore not what these systems cost. The question is who sets when an agent may stop, and whether anyone in your company has ever done it.

Image generated with ChatGPT
Share this article

More article

Article, Podcast
9. September 2026
A pilot is, by definition, the best of several attempts. That is no accusation of anyone, it is the nature of the thing. You…
Article, Podcast
26. August 2026
Forty years ago we began to digitize the world into zeros and ones. To make the world discrete, in the scientific sense. Now we…
Article, Podcast
12. August 2026
Before an agent can act for a company, someone has to write down what a customer is. The sentence sounds like paperwork. It is…
Article, Podcast
9. September 2026
A pilot is, by definition, the best of several attempts. That is no accusation of anyone, it is the nature of the thing. You…
Article, Podcast
26. August 2026
Forty years ago we began to digitize the world into zeros and ones. To make the world discrete, in the scientific sense. Now we…