Almost every organisation we speak to has now seen an impressive AI demo. Someone pasted a document into a chat window, asked a hard question, and got a good answer. The demo took twenty minutes to build and it was genuinely convincing. Then the project stalled for eight months.
The reason is consistent. A demo is a single sample. A system is a distribution. The demo used the document someone chose; production uses the document that arrives on a Tuesday afternoon, scanned at an angle, in a template last updated in 2013, with three contradictory figures in it. Nothing about the demo tells you how the system behaves there.
Build the measurement before the feature
The single highest-leverage thing you can do on an AI project is to build the evaluation harness before you build the product. Take two or three hundred real historical cases where you already know the correct outcome. Score the system against them. Now every subsequent change (a new prompt, a new retrieval strategy, a new model) produces a number rather than an argument.
This sounds like process overhead. In practice it is the fastest route to shipping, because it collapses the endless subjective debate about whether the output is good. It is either better on the set or it is not.
What we measure
- Extraction accuracy against a hand-labelled hold-out set, reported per field rather than as one flattering average.
- Fabrication rate: how often the system asserts something the source documents do not support.
- Refusal quality: whether it declines when it should, and escalates cleanly to a person.
- Cost and latency per completed task, because unit economics decide whether a pilot ever becomes production.
Grounding beats cleverness
The systems that survive contact with real users are the ones that cite their sources. Not because citation is a nice feature, but because it changes the relationship: a reviewer can open the referenced page and confirm the claim in four seconds. Trust is built through verifiability, not through confidence in the tone of the answer.
An AI feature that cannot be audited will be quietly abandoned by the experts it was built for.
We have watched this happen. A well-performing model with no citation trail gets used enthusiastically for three weeks, then progressively less, because senior staff cannot defend a decision they cannot trace. The model was fine. The system around it was not.
Decide where the human sits
The most important design decision in an applied AI project is not the model. It is where a human being enters the loop, and what authority they hold when they do. In the underwriting work we did earlier this year, the model drafts and a person decides; nothing is issued automatically. That constraint did not weaken the business case. Turnaround dropped from six and a half hours to under an hour, because the clerical work disappeared while the judgement stayed exactly where it belonged.
The interesting question is rarely whether a model can do the whole job. It is which part of the job it should do, and how the handoff is designed. Get that right and the technology becomes unremarkable in the best possible way: it simply makes the work shorter.
A practical starting sequence
- Pick one process with high volume, unstructured input, and an expert bottleneck.
- Assemble two hundred historical cases with known outcomes before writing any prompt.
- Build the narrowest possible version that produces a reviewable draft with citations.
- Pilot with the experts who do the work today, and treat every correction as data.
- Only then discuss scale, and price the whole thing per completed task.
None of this is exotic. It is the same engineering discipline that makes any other system trustworthy, applied to a component that happens to be probabilistic. The organisations getting real value from AI right now are not the ones with the cleverest prompts. They are the ones who took evaluation seriously in month one.
About the author
Angel Maile
Angel co-founded Bonang Technologies after a decade spent building software inside organisations where the technology decisions and the commercial ones were made in separate rooms. The company exists to close that gap: engineering that starts from what the business is actually trying to achieve, and leadership that can hold both conversations at once.