Insight
Why 95% of AI pilots never reach production (and what the 5% do differently)
By Brett Raven, APAC Lead & Board Member, AIREUpdated

Three separate 2026 studies, using three different methods, all land on the same place.
What changes
Most AI pilots stall before they ever reach production. MIT's NANDA initiative found 95% of generative AI pilots show no measurable profit-and-loss impact (MIT NANDA, 2025). RAND Corporation analysed over 2,400 enterprise AI initiatives (RAND Corporation, cited in Pertama Partners, 2026). It found 80% fail to deliver their intended business value, twice the failure rate of a typical IT project. HCLTech surveyed 467 executives at billion-dollar-plus enterprises. It found 43% of AI initiatives are at risk, driven by execution gaps rather than technology gaps.
We get you AI-working, not just AI-licensed. So we read these numbers as a warning about how AI programmes are run, not a verdict on the technology. If you lead a mid-market business in Australia, you're weighing where to put your next AI dollar. Here's what the data says, and what to do about it.
- Three independent studies (MIT, RAND, HCLTech) converge on one finding: most AI pilots fail to reach or sustain production. The cause is organisational, not technical.
- Buying AI capability from a specialised vendor succeeds roughly twice as often as building it internally.
- The single biggest predictor of success is whether the workflow was redesigned around the tool, not the other way around.
- Confusing "agent" (a task) with "agentic" (a system) leads teams to govern and measure the wrong thing.
- You can pressure-test your own readiness in one afternoon, before you spend on a pilot.
What three different studies agree on
MIT, RAND, and HCLTech didn't study the same companies, use the same sample, or ask the same questions. That's what makes the overlap worth noting. When three independent methodologies agree, you're looking at a real signal, not a single flawed survey.
Independent failure-rate estimates (2025-2026)
- MIT NANDA (no P&L impact)
- 95%
- RAND (fails to deliver value)
- 80%
- HCLTech (at risk of failure)
- 43%
Different methodologies, different samples, same direction. Not directly comparable figures.
The common thread across all three studies is not model quality. It's what happens after the model gets deployed. MIT's researchers point to a "learning gap." Both the tools and the organisations around them lack the operating discipline to turn a demo into a habit.
This is the same read we give leaders in our advisory work: the technology question comes second, after someone has actually looked at how the organisation runs today.
What's actually killing these pilots

A pilot succeeding in a demo and a pilot succeeding in production are different tests. The demo runs on clean data, a small team, and full executive attention. Production runs on messy data, hundreds of users, and whatever else is competing for their time that week.
Four patterns show up again and again in why pilots die before they scale:
- Licences without habits. Rolling out seats and hoping adoption follows, with no training and no measurement.
- Tool-first thinking. Picking the platform before the problem is even defined.
- The plug-and-play fallacy. Assuming AI drops into an existing workflow untouched, when the workflow is usually the thing that needs to change.
- Agent and agentic, confused. An agent is a single building block: software plus a model, given a goal and a task. Agentic is the system: multiple agents and people handing work back and forth toward an outcome. Teams that conflate the two govern a system like a tool, and measure a tool like a system.
None of these four patterns are technology problems. They're operating-model problems, and they're fixable with discipline rather than budget.
Why "our strategy is AI" is a warning sign
If you hear a leader say their strategy is AI, that's a tell they don't have one. AI is the implementation layer. Strategy is what the business is trying to win at, and that hasn't changed. The question was never "should we experiment with AI." It's "what operating model do we need now that AI can do real work."
We've sat across the table from mid-market leaders who could describe their AI tool stack in detail. They couldn't name the three workflows it was meant to improve. That gap, between tool fluency and workflow clarity, is where most AI budget quietly disappears.
The buy-versus-build signal nobody's talking about
Here's a finding that deserves more attention than it's getting. MIT's research found that buying AI capability from a specialised vendor succeeds roughly twice as often as building it internally.
Success rate by sourcing model
- Bought from a vendor
- 67%
- Built internally
- ~22%
This doesn't mean every business should outsource its entire AI programme. It means the businesses succeeding tend to pair an experienced outside partner with internal ownership of the workflow. They don't ask one internal team to invent the capability and the discipline to run it, all at once.
It's the same split we work to on the build side: platform implementation and custom applications for the process no vendor sells, handed over working rather than left half-finished.
The contrarian read: it's not failing, it's under-resourced
Wharton's Peter Cappelli offers a useful counterpoint to the doom narrative building around these statistics. His argument: AI adoption isn't failing so much as it's being sold as easier and cheaper than it actually is. Doing this properly takes real investment in training and change management, not just a licence and a launch email.
That reading fits the pattern in the data. The 5% of pilots that do scale aren't the ones with the best model. They're the ones that treated adoption as a resourced, measured, ongoing programme rather than a one-off rollout.
What to do before your next AI pilot

Three steps, in order, before any tool gets selected:
- Pick your capabilities. Start with the two or three capabilities most critical to your strategy, not the tools that look exciting this quarter. Everything else waits.
- Map the workflows. Document how the work actually happens today. Look for friction, delay, rework, and manual effort. That's where AI earns its keep, not in a generic use-case list.
- Prioritise and design. Pick the top three opportunities, not the top thirty. Design the solution properly: a named owner, a governance step, and a success metric, before you select any vendor.
If this resonates, AIRE runs a facilitated half-day working session with your senior leadership team. You'll walk out with 3 to 5 priority capability areas and a 30/60/90-day plan, owners named, metrics defined, reviewable in 30 days.
What changes from here
Most failed AI initiatives aren't a tooling problem. They're an operating-model problem instead. The fix looks less like buying more software and more like ordinary operational discipline. Pick the right capabilities. Map the real workflow. Design before you buy.
Before you spend on the next pilot, run our AI Readiness Assessment and see exactly where your organisation actually stands.
Frequently asked questions
Is a 95% failure rate too pessimistic to be useful?
No. The figure comes from MIT's NANDA initiative measuring profit-and-loss impact specifically, a high bar. RAND's independent 80% figure, using a different method, points the same direction. Most AI initiatives underperform their stated goal, even short of full profit-and-loss impact.
Does this mean small and mid-market businesses should wait before investing in AI?
No. It means the order of operations matters more than the timing. Map your workflows first, then pick two or three priority capabilities before selecting a tool. That order beats buying licences first and figuring out the use case later.
What's the difference between an AI agent and an agentic workflow?
An agent is a single building block: software plus a model, given a goal, tools, and guardrails. It does one discrete task, such as drafting a reply or enriching a record. An agentic workflow is the system: multiple agents and people handing work back and forth toward an outcome. Treating the two as the same thing is a common source of confused governance.
Sources
We read each source on the date shown beside it. Studies are revised and re-reported, so check the original before you quote a figure.
- Fortune, MIT report: 95% of generative AI pilots at companies are failing (reporting MIT NANDA, The GenAI Divide: State of AI in Business 2025). Read 10 September 2026.
- Pertama Partners, AI project failure statistics 2026 (citing RAND Corporation). Read 10 September 2026.