I had a small document capture project where a model read handwritten index cards into structured data. Before trusting any of it, I measured how often it was right. The result was not that the model was bad. The result was that the part of it I could have checked most easily told me nothing.
What the measurement showed
Three things were read off each card: its structure, its dates, and its amounts. Structure was reliable. Dates were reliable. Amounts were not.
That much is unremarkable. Handwritten numerals are hard, and anyone who has looked at a stack of real paperwork would guess as much. The part worth knowing is the shape of the failure. Every wrong amount came back with high self-reported confidence. Not most of them. The wrong values did not look any less certain than the right ones, so asking the model how sure it was gave no warning at all.
There was a second check running alongside it: read the same card several times and see whether the answers agree. That one worked. The amounts that came back differently across repeated reads were, overwhelmingly, the ones that were wrong.
Why self-reported confidence cannot do this job
If you have ever worked anywhere near a calibration lab, the reason will be obvious. A model's stated confidence is the instrument grading its own accuracy. It is a number the thing under test produces about itself, using the same process that produced the answer you are trying to check. When the process goes wrong, it goes wrong in both places at once, which is exactly why the confidence score stayed high on the bad reads.
Agreement across repeated reads is a different kind of number. It is repeatability, and repeatability is a property of the measurement system rather than a claim the system makes about itself. Manufacturing has been separating those two ideas for fifty years under the name gage repeatability and reproducibility. You do not ask the caliper whether it trusts itself. You measure the same part several times and look at the spread.
The version of this with no AI in it
The same failure shows up without a model anywhere near it, and that version is more common. An automated step appended a row to a spreadsheet. Every total kept summing the original range. One value sat outside the total from then on, the sheet under-reported, and nothing raised an error at any point.
Wrong, quiet, and confident. A model reporting certainty on a misread number and a spreadsheet reporting a total that silently excludes a row are the same failure from two directions: the output carries no signal that it is wrong, so the only protection is a check that lives outside it.
What this changes about how the work is ordered
Most adoption conversations are about capability. Can it read the cards, can it draft the summary, can it handle the edge case. Capability is usually not the bottleneck. The bottleneck is that nobody has written down what a correct answer looks like, so there is nothing to evaluate against except how good the output feels.
Three things follow from the measurement above, and none of them require any particular tool:
Set the threshold before you see the results. A number chosen after the fact is a number chosen to make the result acceptable. Decide what accuracy would be good enough to act on while you still have no idea whether you will hit it.
Measure per field, not per document. The project above would have looked fine as a single aggregate score, because structure and dates carried it. The useful answer was that two of three fields could be trusted and the third could not, which is an operational instruction rather than a grade.
Use repeated reads, not stated confidence. It costs a little more to run the same input several times, and it is the only one of the two that told me anything.
The honest conclusion from that project was not that the pipeline worked or that it did not. It was that a person should still check the totals, and that everything else could be left alone. That is a smaller claim than either side of the usual argument, and it is the kind of claim you can only make if you measured.
I wrote up the questions worth answering before this kind of work starts as a single page, the evaluation plan, which is free and has nothing to sign up for. The longer argument about why evaluation rather than capability is the bottleneck is in the adoption essay.