There is a particular kind of meeting that has become common over the last two years. A team demonstrates something genuinely impressive — an assistant that answers questions about the company's own documents, a system that reads invoices, a model that drafts responses to customers. Everyone in the room agrees it is remarkable. Then it is never heard of again.
The instinct is to blame the technology. In our experience that is almost never the reason. The model was fine. What was missing was everything around it.
The demo is the easy fifteen percent
A pilot is built against clean inputs, a cooperative audience and the happy path. Production is the opposite of all three. It runs against the documents your business actually has, for users who did not ask for it, and it fails in ways nobody anticipated.
The work that decides whether something ships is not the work that makes a good demo:
- What happens when the system is not confident, and who it escalates to
- How a wrong answer is caught, corrected and prevented from recurring
- Which system of record it writes to, and what happens when that write fails
- Who owns it in six months when the underlying model changes
- Whether the documents it answers from are current, complete and permitted
Data readiness is the real project
A striking share of AI engagements turn out to be information architecture engagements wearing a different name. If the knowledge a system answers from is spread across four shared drives in three formats with six years of superseded versions, no model will save you. It will answer confidently from the outdated document, which is worse than not answering at all.
The unglamorous work of organising what you already know is usually the project. The AI is the interface to it.
We now assess document and data readiness before committing to any deployment, and we are direct when the answer is that the estate needs work first. It is a less exciting conversation, and it is the difference between shipping and not.
No measure, no mandate
Pilots that cannot demonstrate value do not get rolled out, because nobody can make the argument for the budget. This sounds obvious and is routinely skipped. If you cannot say what the system is expected to improve and by how much, you have built a demonstration rather than a capability.
Define the measure before you build. Response time, hours returned, error rate, resolution rate — something countable, agreed with the people who will be asked to fund the rollout.
Put it where the work happens
The last common failure is placement. A capability that lives in its own application asks people to change their habits in order to use it. Most will not. The same capability inside the CRM, the support desk or the document system they already have open gets used without anyone being asked to adopt anything.
None of this is an argument against AI in business. It is an argument for treating it as an operational capability rather than as an experiment — which means scoping the integration, the edge cases and the ownership at the start, not discovering them after the demo went well.