Trust but Verify: Rethinking AI Risk from Pilot to Production
In this blog
In brief
Nearly every enterprise has proven that generative AI works in a demo or proof of concept situation. However, few have built the discipline to scale these to production. MIT's 2025 research found that over 95% of enterprise GenAI pilots deliver no measurable impact to the P&L, and the cause is rarely the model itself but the lack of a good evaluation framework. How can you be sure, without a doubt, that what will be deployed in production will not result in catastrophe? As GenAI scales, what will set leaders apart will be their ability to take the solution to a nagging and costly business problem all the way from a simple POC to a scalable solution deployed in a production environment.
This piece explains why AI risk compounds rather than adds and offers one operating rule: match your evaluation rigor to the stakes of the decision.
The demo works. That was never the hard part
Two years into the enterprise GenAI era, the scoreboard looks good. Boards asked for AI strategies and got them. Budgets have moved. Every function has a pilot underway, and the technology has held up its end of the bargain: today's frontier models are capable, increasingly interchangeable and priced so that access is no longer an advantage. The demo almost always impresses.
Far fewer have turned those pilots into production capabilities they can stake a decision on. MIT's 2025 study, The GenAI Divide, bluntly captured the gap: by its measure, only about 5% of enterprise GenAI initiatives produced rapid, measurable value, while the large majority stalled with little impact on the bottom line.
The most important finding is why the pilots stall. MIT attributes it not to model quality but to what it calls a "learning gap": the failure to wire AI into workflows, structures and, above all, evaluation. The models are good enough. What is missing is the discipline around them. That changes the question executives should be asking. The issue is no longer whether the model can do the work, but how anyone knows the output is right, and what happens when it isn't.
When it fails, nothing breaks
Part of the answer is that generative AI fails in an unfamiliar way. Most of what enterprises use it for is knowledge work: research summaries, market analyses, draft proposals, contract reviews and board memos. When a human analyst gets something wrong, the work usually shows it somewhere. When a model gets something wrong, the output arrives just as fluent, confident and well formatted as when it is right. A citation that leads nowhere. A market-size figure that is plausible and invented. A summary that quietly drops the one clause that mattered. Nothing announces the failure; it sits in a deliverable that already looks finished.
This is how capable organizations end up in the 95%. Every enterprise already has quality control for knowledge work, but it was built for human authors and it runs on proxies: polish signals care, confidence signals competence and a reviewer skims until a sloppy patch says look closer. Generative AI breaks all three. Its output is evenly polished, whether it is right or wrong, and it produces drafts faster than any review chain was designed to absorb. The companies stalling at the pilot stage are not careless. They are discovering that instincts their reviewers have honed over whole careers do not work on this material.
Why AI risk is multiplicative, not additive
In 1993, the economist Michael Kremer, who later won a Nobel prize, published his O-Ring Theory, named for the single brittle seal that destroyed the Space Shuttle Challenger. His insight was that in a system of interdependent tasks, quality multiplies rather than adds. One weak component drags down the value of everything else, no matter how strong the other parts are.
A production GenAI system is exactly that kind of chain: data, model, prompt, retrieval, guardrails and the human who acts on the output. If each link is 90% reliable, six links in sequence leave you near 53%, not 90%. Reliability is the product of the parts, not their average. A single unverified claim, a missing guardrail or an over-trusting user can bring down an otherwise excellent system.
The Challenger analogy carries a second, subtler warning: normalization of deviance. Engineers had seen O-ring erosion on 24 earlier flights, but because none of them ended in disaster, the damage was quietly reclassified from "anomaly" to "acceptable risk." The same drift happens with AI. Every time a plausible but unverified output ships without incident, the organization quietly resets what "good enough" means. Caution keeps giving ground to speed until the first real failure arrives at scale.
Match your evaluation rigor to the stakes
If risk multiplies, the answer is not to verify everything to the same degree, which is slow and expensive, nor to trust everything equally, which is how most AI failures happen. The discipline is to stack verification methods in proportion to the stakes of the output. A throwaway brainstorm needs almost no scrutiny. A client-facing deliverable earns the full stack.
Five methods, layered from cheapest to most rigorous:
- Self-review: Flag claims that feel off: logical jumps, suspiciously round numbers, confident specifics.
- AI probe: Push back on the same response: ask for sources, ask for the counterargument, ask "what would change your answer?"
- Source triangulation: Check claims against authoritative material: official statistics, primary research, a database or a different model run on the same prompt.
- Expert / SME validation: Route the output through a domain expert who can spot both errors and, just as important, what was silently left out.
- Empirical / real-world test: Back-test against historical data, and pilot on a single business unit before scaling.
The differentiating skill is managerial, not technical
Working with a capable model has less in common with running software than with managing a bright, fast and overconfident new analyst. The skills that separate leaders are the ones good managers already have: explaining what they need with clear context, giving feedback and iterating and defining what "done" looks like with checkpoints that verify quality. Structured thinking, stakeholder management and quality control are governance muscles an enterprise already has, and they are exactly what AI systems need wrapped around them.
This is what makes "trust, but verify" an operating model rather than a slogan. Access to good models is no longer the differentiator, since everyone has it. The organizations pulling ahead are the ones that have made disciplined evaluation a habit, so they can move fast and still trust what they ship.