Why AI Pilots Fail: What the 5% Did Differently

MIT found roughly 95% of enterprise AI pilots show no measurable return. What the 5% did differently: where it was aimed, who built it, and what was measured first.

· Mahdy Hasan · AI & ML

MIT's Project NANDA found roughly 95 percent of enterprise generative AI pilots produced no measurable P&L impact, despite 88 percent organisational adoption and $581.7 billion in corporate AI investment. The published data points at three splits rather than model quality. Budgets went to sales and marketing while documented returns appeared in back-office automation. Vendor-built tools reached production about 67 percent of the time against 33 percent for internal builds. And pilots launched without a recorded baseline cannot prove a return even when one exists.

AI pilots fail mostly because the system around the model is missing. MIT's Project NANDA found roughly 95 percent of generative AI pilots showed no measurable P&L impact, citing brittle workflows and poor fit with daily operations. Budgets concentrated in sales and marketing, while the clearest documented returns appeared in back-office automation.

A client asked me a version of this question last quarter. They had shipped an AI feature, the team liked it, and the finance director wanted to know what it returned. Nobody could answer.

That situation is close to the norm. The published 2026 data explains why, and it points somewhere more useful than model quality.

  • MIT's Project NANDA found roughly 95 percent of enterprise AI pilots produced no measurable P&L impact.
  • Vendor-built AI tools reached production about 67 percent of the time, against 33 percent for internal builds.
  • More than half of generative AI budgets went to sales and marketing, while documented returns appeared in back-office automation.
  • Organisational AI adoption reached 88 percent and corporate investment hit $581.7 billion, up 130 percent (Stanford AI Index 2026).
  • Workers aged 22 to 25 in AI-exposed occupations show a 16 percent relative employment decline, from reduced hiring rather than layoffs.
  • Employee AI use exceeds 80 percent in India, China, Nigeria, the UAE, Egypt, and Saudi Arabia. The United States ranks 24th on population adoption.

Why Do 95% of AI Pilots Fail to Show Returns?

AI pilots fail because the system around the model is missing, not because the model is weak. MIT's Project NANDA reviewed enterprise deployments and found roughly 95 percent produced no measurable effect on profit and loss.

A caveat belongs next to that number. The MIT work is preliminary and not peer reviewed. It defines success narrowly, as documented P&L impact inside about six months. Real value often fails that test.

The figure is still worth using. It matches the pattern I see when clients bring us a stalled pilot to review.

The GenAI Divide

The GenAI Divide is the term MIT's Project NANDA uses for the gap between widespread generative AI adoption and the small share of deployments producing measurable business returns. In its 2025 study, about 5 percent of pilots delivered documented P&L impact while the rest showed adoption without transformation. The report attributes the divide to a learning gap in tools and organisations rather than to model quality.

95% Of enterprise generative AI pilots showed no measurable P&L impact MIT Project NANDA, The GenAI Divide: State of AI in Business 2025

What Did the 5 Percent Do Differently?

They aimed at operations work, bought more than they built, and measured a baseline first. Two of those three splits come straight out of the MIT findings.

The first split is where the work was pointed. More than half of generative AI budgets went to visible top-line functions like sales and marketing. The documented returns showed up in back-office automation. They often arrived through reduced external spend, such as cut agency fees and replaced BPO contracts.

The second split is who built the thing. Tools bought from specialist vendors reached production about 67 percent of the time. Internally built tools managed about 33 percent, roughly half as often.

LATECH5, Kuala Lumpur

An autonomous operations platform for Malaysian SMEs, built on the unglamorous side of the ledger: customer service, HR compliance, inventory routing, and finance workflows, running inside WhatsApp and Telegram where those businesses already work. Augmex delivered the beta in five months with a six-person team, including an NLP layer handling code-switching across Bahasa Malaysia, English, Mandarin, and Tamil. The platform was ready to onboard its first 50 SMEs on schedule.

Read the full case study

When a client asks why their pilot stalled, my first question is never about the model. I ask what number they wrote down before they started. Most cannot answer, and that is usually the whole diagnosis.

Mahdy Hasan, Founder & CEO, Augmex

How Do You Tell If Your Own Pilot Is Failing?

Answer six questions about the pilot you already have. Missing answers, rather than bad ones, are the reliable warning sign.

  1. What number was recorded before launch? Hours on a task, error rate, or tickets resolved. Without it, no return can be proven later.
  2. Which single workflow does this serve? A tool serving four teams loosely usually serves none of them measurably.
  3. Who owns the output quality? Name one person. A committee means nobody reviews the failures.
  4. What is the cost per successful outcome? Not per API call, and not per seat. Successful means the user kept the result.
  5. What happens when it is wrong? Undocumented escalation paths are where quiet trust damage accumulates.
  6. Could this have been bought? If a vendor tool covers 80 percent of the need, the build has to justify the other 20 percent.

Question one carries most of the weight. A pilot without a baseline gets counted as a failure even when it worked. That inflates the reported failure rate.

What Does the Hiring Data Say About Your Team Plan?

The measurable employment effect so far sits on early-career workers, and it came from hiring decisions rather than layoffs. Brynjolfsson, Chandar, and Chen at Stanford's Digital Economy Lab tracked payroll records using ADP data.

Workers aged 22 to 25 in the most AI-exposed occupations show a 16 percent relative employment decline. That figure controls for firm-level shocks. For software developers in that age band, the drop runs about 20 percent from the late-2022 peak. Older workers in the same occupations held steady.

The mechanism matters for planning. Firms did not cut juniors. They stopped opening the roles, which is a quieter decision and easier to repeat by default.

That leaves a question worth putting on a hiring plan now. The mid-level engineers you need in 2029 have to come from somewhere. We keep training juniors at Augmex partly for that reason, and the spreadsheet argues against it most quarters.

Which Countries Actually Use AI the Most?

The countries building the frontier models are not the countries using them hardest. Stanford's 2026 AI Index puts employee AI use above 80 percent in six countries. Those are India, China, Nigeria, the UAE, Egypt, and Saudi Arabia. The global figure is 58 percent.

24th United States rank on population AI adoption, at 28.3 percent. Singapore leads at 61 percent Stanford HAI, 2026 AI Index Report

I see the reason from where I sit. Teams here adopt quickly because experiments are cheap to run. The alternative is often hiring nobody at all.

This changes who you are comparing when you plan a delivery model. A team in Dhaka, Lagos, or Bangalore running AI-assisted delivery is no longer competing on rate alone.

Why Is Verification Becoming the Expensive Part?

Verification is becoming expensive because only half the work got cheaper. Producing a draft, a contract review, or a first-pass analysis now costs close to nothing. Confirming any of it is correct costs roughly what it did in 2023.

The same shift moved the engineering constraint from writing code to designing systems. I covered that separately in a piece on AI-native architecture.

Agent reliability sharpens the point. Stanford reports agent success on real-world tasks moving from 20 percent to 77.3 percent in a single year. Budget for the checking, because that is the half that stays expensive.

If your pilot sits in the 95 percent, the fix is rarely a better model. It is a named workflow, a recorded baseline, and one person who owns the output. That is the conversation our team has most weeks. We will say so plainly when the honest answer is fixing a process, not buying software.

Related Resources

Related Articles