The override log
An agent that never gets overruled is not good. It is unmeasured. The moment somebody who knows the domain says no, not that one, and here is why, you have produced something no model release will ever hand you.
Almost everyone running agents logs what the agent did. Very few log what they did about it, and that second log is worth more than the first.
Think about what is actually happening in that moment. Somebody who knows the domain looked at a confident recommendation, on a specific case, with real stakes attached, and said no. Not that one. Here is why. And then something happened afterward that either proved them right or did not.
There is no other way to manufacture that. It is expert judgment, in writing, attached to a concrete situation, with a checkable outcome on the end of it. You cannot buy it and no lab can ship it to you, because it is not knowledge about the world. It is knowledge about your world.
Why it gets thrown away
Because it feels like a failure. The system got it wrong, you fixed it, the work went out, and the fix reads like cleanup rather than output. So it lives in a chat message or a mental note and evaporates by Thursday.
It is the opposite of cleanup. It is the highest-signal thing that happened all day, and it arrived free, as a byproduct of work you were doing anyway.
The correction is not the cost of running the system. It is the product of running the system.
What to actually capture
Three things, and the third is the one people skip. The diff, meaning what you actually changed, kept verbatim rather than summarized. The reason, in your own words, because "wrong" is not a reason and will teach nothing. And the outcome, whenever it lands, which is often weeks later and is the only part that can tell you whether the override itself was right.
Summarizing the diff is the most common way to ruin this. The specific words a person chose when they rewrote something are the signal. Paraphrase them and you have kept the fact that a correction happened while throwing away what it taught.
The number that should worry you is zero
An agent nobody overrules is not a good agent. It is an unmeasured one. Zero overrides means one of three things, and all three are bad: nobody is really reading the output, the gate is theater, or the work is so low-stakes that being wrong costs nothing.
I would rather see a high override rate that is falling than a low one that has been flat since the day it was set up. Falling means the corrections are landing. Flat means nothing is being learned by anybody, including me.
What it buys
First, better scoring, which is the obvious one. If the last several dozen recommendations of a certain shape all got waved through, that shape has earned some confidence. If a different shape gets rewritten every time, it has earned a demotion.
Second, and this is the part that changes how a day feels, it buys permission. A system that can see its own track record on a category of decision can start telling the difference between a call I would want to make personally and one I would rubber-stamp. That is not a model capability. It falls out of the log.
The permission follows the evidence, never the other way around.
Honest status
Every override I make gets captured. The scoring reads them. What I cannot yet show you is a clean attribution from a specific correction to a specific outcome at any real volume, because that needs more closed cases than I have. Until then the weights are informed hypotheses rather than measured facts, and I would rather say that than dress it up.
This is the thing underneath Evidence compounds. It is also the mechanism behind what happens when the loop closes, and the grading it feeds is in How I grade my agents.
