The evaluation plan
One page to fill in before you build. It is the document that lets a careful reviewer say yes, and the reason most artificial intelligence pilots never get one.
Why this page exists
The pattern is consistent. A team builds something impressive. Everyone agrees it is impressive. Then it has to touch a real decision, and somebody whose job is to be careful asks how we will know when it is wrong. There is no written answer, so they decline, and they are right to.
That is not a governance problem or a culture problem. It is a missing artifact. This is the artifact. It takes about twenty minutes and it is the cheapest thing you will do all quarter, because the alternative is discovering the same gap after the build.
The argument behind it is written up at length in sequencing AI adoption in a regulated organization. You do not need to read that first. This page stands on its own.
The six questions
Answer them in plain language, in a shared document, with the person who owns the process in the room. Not in a ticket, and not alone.
1. What must this produce?
One or two sentences describing a thing, not an intention. "A first-draft response to a supplier query, for a human to approve" is a specification. "Improve supplier response times" is a hope, and nobody can evaluate a hope.
2. How will we know it is right?
This is the one that carries everything, and the one almost everyone skips because it feels obvious until you write it down. Describe an input and the exact output it should give, so that somebody else could check it without asking you.
If the honest answer is that a person would have to read every output and use judgment, write that down too. It is still an answer, and it tells you the real running cost before you commit to it rather than after.
If you cannot answer this one, stop. That is not a delay, it is the finding. An unwritten standard is indistinguishable from no standard, and a reviewer with no standard will correctly decline. Build nothing until this question has an answer.
3. What must it never do?
One-line rules, usually about money, personal data, or anything irreversible. Never invent a figure. Never guess at a date. Never send anything outward without a person seeing it first.
These go unwritten because they feel too obvious to say. They are not obvious at all to a system that will otherwise fill the gap with something plausible.
4. What counts as good enough to widen?
A number, and the sample it is measured on, written down before you see any results. Deciding what counts as success after you know the score is how careful organizations talk themselves into shipping things they should not have.
Include the number that actually predicts safety, which is not accuracy. It is how often the system is wrong while looking right. A confident, well-formatted, subtly wrong answer is the expensive failure, because nothing raises an error and it is found a month later by somebody doing an audit.
5. Who signs this off, and what do they need to see?
Name the person or the committee, and ask them now what evidence they would need. This is the question that separates an organization where approval takes a week from one where it takes two quarters, and asking it early costs nothing.
Most reviewers are not obstructive. They are being asked to accept a risk nobody has described to them in terms they can weigh.
6. What is explicitly out of scope?
Write down what this is not, or it will grow while you are not looking. This is the line that protects a fixed timeline, and the one most often left out.
Two things that make it work
Keep it to one page. The moment it becomes a document with an owner and a review cycle, it stops being the thing that unblocks the build and becomes another thing to get through. If it does not fit on a page, the scope is too big, which is itself useful to know.
Read the edits. Once people are reviewing output, every correction they make is a requirement you failed to write down. Not a correction, a specification. Collect them and they become the acceptance criteria you were missing, and review stops being a permanent tax and starts converging.
If this was useful
I write most Sundays about adopting this technology where being wrong is expensive: what worked, what failed quietly, and the measurements behind both. No pitch, and you can leave whenever you like.
If you would rather not do it alone
I take on a small number of fixed-scope engagements that produce exactly this, and the working prototype that proves it. The consulting page has what they are and what they cost, and I do not take work in the financial sector.
If you think one of these questions is wrong, or that there is a seventh, I would genuinely rather hear that than not. The contact page works.