DevXen
All insightsAI strategy5 minute read

The 95% Statistic: What the MIT GenAI Report Actually Says

The MIT Project NANDA report is quoted constantly and read rarely. Here is what it measured, where it contradicts itself, and what it justifies.

"95% of AI projects fail" is now the most quoted statistic in enterprise software. It is quoted by vendors selling the cure and by sceptics calling the whole field a bubble. Both are misreading the source.

The number comes from one specific report. Reading it properly changes what you should do next.

What the source actually is

The report is The GenAI Divide: State of AI in Business 2025, published in July 2025 by MIT Project NANDA. Authors: Aditya Challapally, Chris Pease, Ramesh Raskar, and Pradyumna Chari.

Its stated methodology covers January to June 2025 and combines three things: a systematic review of over 300 publicly disclosed AI initiatives, structured interviews with representatives from 52 organisations, and survey responses from 153 senior leaders collected across four industry conferences.

It is a preliminary research report, not a peer-reviewed study, and the authors state that the views are their own and do not reflect the positions of affiliated employers.

What the 95% figure measures

From the executive summary: despite 30 to 40 billion dollars of enterprise investment in generative AI, "95% of organizations are getting zero return," with just 5% of integrated pilots extracting millions in value while "the vast majority remain stuck with no measurable P&L impact."

Three qualifications matter.

It measures generative AI specifically, not machine learning, forecasting, workflow automation, or the rules-based integration work that quietly runs most operations.

It measures measurable profit and loss impact, not usefulness. The report notes separately that general tools like ChatGPT and Copilot are widely adopted and do improve individual productivity. They just do not show up in the accounts.

It measures custom and vendor-sold enterprise deployments. The pipeline reported for task-specific tools is 60% of organisations evaluating, 20% reaching pilot, and 5% reaching production. General-purpose chatbots showed a pilot-to-implementation rate of roughly 83%.

So the accurate sentence is: in the deployments this report examined, about 95% of custom enterprise generative AI efforts produced no measurable P&L impact. That is not the same as "95% of AI fails."

The finding that matters more than the headline

The report is explicit that the barrier is not what most people assume. Listed among five myths it sets out to correct:

The biggest thing holding back AI is model quality, legal, data, risk. What's really holding it back is that most AI tools don't learn and don't integrate well into workflows.

The authors call this the learning gap. Systems that cannot retain feedback, carry context between sessions, or adapt to how a specific business actually operates stall after the pilot, regardless of which model sits underneath.

Four patterns are named in the executive summary:

  • Limited disruption. Only a small number of sectors show meaningful structural change.
  • Enterprise paradox. Large firms lead in pilot volume but lag in scale-up. Mid-market top performers reported roughly 90 days from pilot to full implementation, while enterprises took nine months or longer.
  • Investment bias. Budgets favour visible, top-line functions over back office work with higher return.
  • Implementation advantage. External partnerships were reported to succeed at twice the rate of internal builds.

On investment bias, the report describes executives allocating a hypothetical hundred dollars across functions, with sales and marketing dominating. Its own research note attributes that to attribution convenience rather than value: demo volume and email response time map neatly onto board KPIs, while fewer compliance breaches and a faster month-end do not.

Where the report contradicts itself

Two internal inconsistencies are worth flagging, because they show why the precise figures should be handled carefully.

The executive summary says "only 2 of 8 major sectors show meaningful structural change." The body of the report says seven of nine sectors show little to no structural change. Those are different denominators.

On budget allocation, the body text says sales and marketing captured "approximately 70 percent" of AI budget allocation, while the accompanying research note says "~50% to Sales & Marketing." Both appear in the same document.

The authors are more candid than most of the people quoting them. Their own research limitation note reads: "These figures are directionally accurate based on individual interviews rather than official company reporting. Sample sizes vary by category, and success definitions may differ across organizations."

Directionally accurate is the right way to use this report. It is not a measurement instrument.

Independent corroboration

Two Gartner forecasts point the same direction from a different method.

On 25 June 2025, Gartner predicted that more than 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls.

Earlier, on 29 July 2024, Gartner predicted that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025.

These are analyst forecasts rather than measurements. Their value is that three separate exercises, using different methods, converge on the same failure mode: pilots that were never connected to a measurable business outcome.

What we take from this

The report is not evidence that AI does not work. It is evidence about which approach works, and the pattern is consistent with what we see in service businesses.

Pick the process before the tool. The failures described are integration failures. A model evaluated in isolation tells you nothing about whether it survives contact with your approvals, your exceptions, and your data.

Insist on a measurable outcome first. If a deployment cannot name the number it should move, it will join the 95% by definition, because no measurement was ever set up.

Do not read the back office as low value. The report's own conclusion is that back-office automation is underfunded relative to its return, precisely because it is harder to show off.

Be sceptical of internal-build enthusiasm. Internal builds were reported to fail twice as often as external partnerships. We have a commercial interest in that finding, which is exactly why you should check it against the source rather than take our word for it.

The useful version of the statistic is not "95% of AI fails." It is: most generative AI spending goes to systems that never touch a measured business process, and that is a choice made before any technology is selected.

Sources

  • Challapally, A., Pease, C., Raskar, R., and Chari, P. The GenAI Divide: State of AI in Business 2025. MIT Project NANDA, July 2025.
  • Gartner press release, 25 June 2025: "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027."
  • Gartner press release, 29 July 2024: "Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025."