The short answer
MIT's NANDA research found roughly 95% of enterprise generative AI pilots deliver no measurable profit-and-loss impact, and RAND puts the broader AI project failure rate above 80%, about twice that of conventional IT projects. The causes are consistent and none of them are the model: no agreed definition of success, weak data foundations, and no integration into real workflows. Projects with quantified success metrics defined upfront succeed at 54%. Those without succeed at 12%.

Five independent studies from Gartner, MIT, RAND, BCG and McKinsey converged on the same range across 2025 and 2026.
95%
of GenAI pilots deliver no measurable P&L impact (MIT NANDA)
80%+
AI project failure rate, twice that of conventional IT (RAND)
73%
of failed projects had no agreed definition of success
61%
were approved on projected ROI that was never measured after launch
MIT defined successfully implemented strictly: sustained productivity gains and documented P&L impact, verified by both end users and executives. Most deployments are simply not measured to that standard, which means the figure captures value realisation rather than model capability.
That distinction matters, because the common reading of the statistic is that the technology does not deliver. The research says something different and more useful: most organisations underestimate the data governance and engineering rigour required to move from an impressive demo to a reliable production system, and then never check whether the result was worth the money.
Read the list of causes again. Not one item is about the model. When five separate analyses using different methodologies converge on the same conclusion, you are looking at a systemic problem in how projects are run, not an implementation detail.
We want to use AI is not a project. Reduce the time to answer a customer order-status question from four hours to four minutes is a project. The second can be measured; the first can only be declared successful.
Models inherit whatever mess exists underneath. If customer data sits in three systems with different spellings of the same company name, an AI layer produces confident nonsense. The audit belongs before the pilot.
Automating a bad workflow makes it fast and bad. If a process needs six approvals because nobody trusts the data, AI does not fix that, it just produces the untrusted output faster.
Over half of AI budgets in 2025 went into sales and marketing pilots, which are high visibility and low ROI. MIT found the real returns came from back-office automation, which nobody demonstrates at a board meeting.
The shape matters more than the technology. Pick something high-volume, rule-heavy, and currently done by a person who dislikes doing it.
Order status queries. Invoice data entry. First-line support triage. Document classification. All four are repetitive, rule-bound and measurable, which is exactly what makes them suitable.
How long does it take, how often does it go wrong, how many times a week does it happen. Write the numbers down before anything changes, because reconstructing them afterwards is guesswork.
Can a system reach the data it needs through an API or an export, and does that data agree with itself? If not, that is the project, and it is worth doing regardless of whether AI follows.
Not the whole process. The most repetitive step in it. A narrow automation that works beats a broad one that half works.
Same metrics, same method. If the numbers moved, scale. If they did not, stop and say so. Both are acceptable outcomes; carrying on without measuring is not.
Our clearest measurable result came from an e-commerce platform where AI generates the question bank. The number that mattered was time: work that took a full day now does not. That is the shape of a project worth doing, one process with a before and an after you can state plainly.
We have also processed 7.6 million documents and now scrape 200,000 documents a month automatically for a client CRM. That is automation rather than AI, and it is worth being precise about the difference, because a lot of what gets sold as AI is deterministic automation that would work more reliably without a model in the loop.
The scoping question we ask first is not which model. It is which process, measured how, and what happens when the system gets it wrong. Projects that cannot answer those three do not get built.
We scope against a measurable process, capture the baseline before building, and tell you when the honest answer is that automation is the better fit.