Guides

Why most AI pilots never reach production, and what the studies actually measured.

By Zain M · Updated 14 September 2026 · 13 min read

Most AI pilots never reach production because the prototype is a fifth of the work. MIT’s 2025 study found only 5% of custom enterprise AI tools reach production and Gartner predicted at least 30% of generative AI projects would be abandoned after proof of concept. Integration, running cost, ownership and an agreed definition of good enough decide the rest.

Custom AI tools reaching production5% (MIT NANDA, 2025)
Projects abandoned after proof of conceptAt least 30% (Gartner)
UK adopters reporting higher revenue12% (DSIT, 2026)
The prototypeRoughly 20% of the work

A demo and a product are different objects

A demo runs on one laptop, with one person watching, who is forgiving and knows which inputs work. Production runs unattended, on real data, for people who will not tolerate a wrong answer and will not report it either; they will simply stop using it. The demo is the part that is pleasant to build and pleasant to watch, which is why so much budget goes into it. Almost every stalled AI project we see is a good demo that nobody costed the distance from, and the distance is roughly four times the demo.

That distance is now measured. Since we first published this guide, three large studies and two UK government surveys have put numbers on how often the demo becomes the product, and the numbers are worse than most buyers assume. This revision sets them out, with their bases and their limits, and then does the practical part: the seven ways a pilot fails, what to measure before you start one, and a worked plan from pilot to production priced against published figures.

One caution before the statistics. The headline numbers come from studies of large American enterprises, and a twenty-person UK firm is not a Fortune 500 company: it has fewer systems to integrate, less data to clean, and no innovation budget to protect a pilot that is not working. Some of that makes production easier to reach, and some harder. Where the UK data differs, and it does, we say so.

What the published studies actually measured

Five figures are quoted everywhere on this subject, usually without their bases. Here they are with the bases attached, because a 95 per cent failure rate for one thing and a 12 per cent success rate for another are not the same claim. The first two rows describe large, mostly American enterprises and are the source of most of the headlines; the last three are UK government and statutory surveys, and they are the rows a UK buyer should weigh most.

Source and dateHeadline figureWhat it measuresBase
MIT NANDA, Jul 202595% of organisations getting zero return; 5% of custom enterprise AI tools reach productionMeasurable profit and loss impact from generative AI; deployment rate of task-specific tools52 organisation interviews, 153 leaders surveyed, 300+ public initiatives; the report calls its figures "directionally accurate"
Gartner, Jul 2024At least 30% of GenAI projects abandoned after proof of concept by end of 2025A prediction, with four causes: poor data quality, inadequate risk controls, escalating costs, unclear business valueGartner analyst research; the accompanying survey was 822 business leaders, Sep to Nov 2023
DSIT AI Adoption Research, Feb 202675% of adopters report improved productivity; 12% report increased revenue; 77% no change in revenue yetSelf-reported outcomes among UK businesses currently using AI3,500 UK business interviews, 12 Feb to 2 May 2025
PwC 29th UK CEO Survey, Jan 202621% say AI increased revenue in the last 12 months; 74% little or no change in revenue; 60% little or no change in costsUK chief executives’ view of AI’s effect on their own numbersUK sample of PwC’s global CEO survey
ONS, Jul 202635% of businesses with 10+ staff use AI; 10% of adopters use it extensivelyAny AI use, and depth of use, in the Business Insights and Conditions SurveyBusiness Insights and Conditions Survey, businesses with 10 or more employees, June 2026

Read on the source pages on 14 September 2026. The MIT and Gartner figures describe large enterprises and are not UK-specific; the DSIT, PwC and ONS figures are UK.

The pilot-to-production funnel, in MIT’s numbers

The MIT report is the one worth reading in full, because it separates two things the press coverage merged. General-purpose tools such as ChatGPT and Copilot are widely adopted: over 80 per cent of the organisations studied had explored or piloted them and nearly 40 per cent reported deployment. The report’s point is that these tools improve individual productivity, not the profit and loss account.

Enterprise-grade systems, custom or vendor-built, are a different story. Sixty per cent of organisations evaluated such tools, 20 per cent reached pilot stage, and 5 per cent reached production. The reasons the report gives are brittle workflows, a lack of contextual learning, and misalignment with day-to-day operations; it calls the underlying problem a learning gap, tools that do not learn, integrate poorly, or match workflows.

Two further findings matter for a small UK firm deciding who should build. In MIT’s sample, external partnerships with learning-capable, customised tools reached deployment about 67 per cent of the time, against about 33 per cent for internally built tools, a difference the authors say was consistent across interviewees though self-reported. And although half of AI budgets went to sales and marketing, the most dramatic cost savings the researchers documented came from back-office automation, with faster payback and clearer cost reductions.

Prototype works100%Integrated with real systems62%Cost per user modelled41%Owner and support agreed27%Live with real users18%
Where AI projects lose momentum between demo and production

What UK businesses report: productivity yes, revenue not yet

The UK evidence is more encouraging than the American headlines, provided you read the right line. DSIT’s 3,500-business study found that three quarters of adopters, 75 per cent, reported improved workforce productivity. But 77 per cent had not yet seen a change in revenue, and just 12 per cent reported an increase. PwC’s survey of UK chief executives in January 2026 found the same shape at the top of the organisation: 21 per cent said AI had increased revenue in the previous twelve months, 74 per cent reported little or no change in revenue, and 60 per cent little or no change in costs. Only 45 per cent said their organisation had a clearly defined roadmap for AI.

Gartner’s own survey of 822 business leaders, run in late 2023, found early adopters reporting on average a 15.8 per cent revenue increase, 15.2 per cent cost savings and a 22.6 per cent productivity improvement. Those are averages among people who had already got something working, which is the point: the returns exist, they are concentrated in the minority that reached production, and the distribution is what a buyer should plan around.

The ONS adds the depth. Of the 35 per cent of businesses with ten or more staff that use AI, only one in ten uses it extensively. Most UK adoption is a tool in one person’s browser, not a system in a process. A pilot that stays a tool has not failed; it has simply not started the journey this guide is about.

The seven failure modes, with a UK small-company angle

The four blockers in our original guide still bite in the order we gave them. The published studies add three more, and for a small UK company each of the seven has a particular shape that the enterprise literature misses: fewer systems but older ones, less data but messier, and an owner who is usually the founder. Together they are the checklist we run on any stalled pilot, and most stalled pilots fail on at least three of them at once.

01Integration. The pilot used a spreadsheet export. Production needs to read and write the system people actually work in, and that is where the real effort sits. In a small firm the system is often a practice-management tool with no API, which is why this is the first thing to check, not the last.
02Cost. Nobody modelled what it costs per user per month, so the finance conversation arrives after the build rather than before it. Gartner names escalating costs as one of its four causes of abandonment. The model is rarely the expensive line; the messages and minutes around it are, and they scale with use.
03Ownership. The enthusiast who built the pilot has another job. Without a named owner and a support arrangement it decays quietly. In a twenty-person company the owner is usually the founder, who has least time of anyone, so the arrangement has to be designed rather than assumed.
04Definition of good enough. Without an agreed accuracy threshold and a plan for what happens when the model is wrong, there is no moment at which anyone can say it is ready. Gartner’s "unclear business value" is this problem seen from the boardroom.
05Data quality. Gartner’s first-listed cause. A pilot runs on the fifty clean examples somebody chose; production runs on the scanned PDF, the email with the attachment missing and the record with the postcode in the wrong field. Small firms have less data and, usually, messier data, so the sample the pilot is tested on must include the ugly cases.
06Skills and the wrong builder. DSIT found 60 per cent of UK businesses cite limited AI skills as a barrier, and the ONS found 62 per cent of those citing an expertise gap are training their own staff rather than hiring help. Both are sensible. But MIT’s finding that external partnerships deployed about twice as often as internal builds is a warning about who should carry the production stage.
07Risk controls and compliance. Gartner’s "inadequate risk controls". In the UK this means a DPIA where personal data is involved, a lawful basis, a written agreement with the supplier and the model provider, and consent and opt-out handling for anything that contacts people. Retrofitting these after the first complaint costs more than building them in.

What a pilot must measure before it starts

A pilot without a baseline cannot succeed, because nobody can tell whether it did. These are the numbers to write down before the first line of code, and the reason each one matters. They take a week to gather, most of it spent timing the task as it is done today, and they are the same numbers the production business case will need, so nothing is wasted. If the supplier cannot help you fill in this table in the scoping session, that is the answer to whether they should build it.

MeasureSet before the pilotWhy it decides the outcome
Baseline hoursHow long the task takes today, per item and per week, measured not estimatedThe only number the saving can be calculated from
VolumeItems per month now, and the realistic ceilingRunning cost scales with this, not with headcount
Accuracy thresholdThe error rate you can live with, and what happens on an errorDefines "finished"; a drafting tool tolerates more than an unattended one
Test setA sample of real inputs including the messy onesData quality is Gartner’s first cause of abandonment
Integration pathWhich system must be read and written, and whether it has an APIThe largest share of production effort
Run-cost ceilingA monthly cap per user or per item, in poundsTurns cost from a surprise into a decision
OwnerA named person with time allocated after go-livePilots without an owner decay
Stop ruleThe result at which you will not proceedSaves the money most pilots spend after they have already failed

A worked example: from pilot to production on the ladder

An illustration, not a client account. A 25-person distributor keys about 2,000 supplier invoices a month into its accounts package. Timed over a week, each takes about six minutes, so the task is roughly 200 hours a month. At a loaded cost of £15 an hour, an assumption you should replace with your own, that is £3,000 a month, or £36,000 a year. That baseline is the whole business case; everything else is arithmetic.

The pilot is an audit. On our ladder that is £2,500 to £6,000, fixed, one to two weeks, credited against any build. It tests extraction on 200 real invoices including the scanned and the handwritten, agrees an accuracy threshold (say 98 per cent of line totals correct, with everything below a confidence score routed to a person), confirms the accounts package has an API, and prices the running cost. Suppose the audit concludes that two thirds of invoices can be posted without a person touching them.

Production is a single-process build at £4,000 to £10,000. The running cost, at today’s published prices, is small enough to state. Reading 2,000 invoices a month at roughly 3,000 input tokens and 500 output tokens each is 6 million input and 1 million output tokens. On OpenAI’s gpt-5-mini ($0.25 and $2.00 per million) that is $3.50 a month; on Anthropic’s Sonnet 5 ($2 and $10) it is $22; even on the most expensive model on any of the three pricing pages, Anthropic’s Fable 5.1 at $10 and $50, it is $110. Hosting and document storage will cost more than the model.

Against a saving of about £2,000 a month from the two thirds automated, a build at the midpoint of the band pays back in roughly four months, and the year-one total, audit credited, is about £8,000 plus a few hundred pounds to run. The point of the arithmetic is not the answer, which will differ for every business, but that every input to it was known before anything was built, so the decision to proceed was a decision rather than a hope.

What to do if you already have a stalled pilot

Do not restart it. Scope the production version specifically: which system it must integrate with, what it costs per user or per item, what accuracy is acceptable and what happens when it falls short, and who owns it afterwards. Fill in the table above for the pilot you already have; most of the blanks will explain why it stopped, and the exercise takes an afternoon rather than a budget.

Then treat the pilot as what it was, a sales tool that won the decision internally, rather than as a foundation to build on. Rebuilding the working part is usually faster than adapting a prototype that was never meant to carry load. MIT’s 5 per cent are not the organisations with the best demos; they are the ones that treated the demo as the start of the costing rather than the end of it.

Method and sources

Every figure above was read on its source page on 14 September 2026. The MIT NANDA report was read in the published PDF; its authors describe the pilot and production figures as directionally accurate and based on interviews rather than company reporting, and we have quoted them with that caveat. Gartner’s figures are from its 29 July 2024 press release, including the 822-leader survey it cites. DSIT’s figures are from the AI Adoption Research report published 13 February 2026, with fieldwork from February to May 2025. PwC’s are from the UK pages of its 29th CEO Survey, January 2026. ONS figures are from the July 2026 article on AI in UK businesses.

Model prices are from the OpenAI and Anthropic pricing pages on the day of writing and are in US dollars excluding VAT. The worked example uses stated assumptions, marked as such in the text, and our own published price ladder; it is not a client account and the hourly cost in it should be replaced with your own. We will revise this page as the sources change, and the date at the top is the date of the last check.

Common questions

What percentage of AI pilots fail?

It depends what is counted. MIT NANDA’s 2025 study found 95 per cent of organisations getting no measurable return from generative AI and only 5 per cent of custom enterprise tools reaching production. Gartner predicted at least 30 per cent of generative AI projects would be abandoned after proof of concept by the end of 2025. Both describe large enterprises; UK small-firm data is thinner and more mixed.

Why do AI projects fail after a successful proof of concept?

Gartner gives four causes: poor data quality, inadequate risk controls, escalating costs and unclear business value. In our experience the order they bite in a small firm is integration first, then unmodelled running cost, then the absence of an owner, then the lack of an agreed accuracy threshold that would let anyone declare it finished.

Do UK businesses see a return from AI?

Productivity, mostly; revenue, rarely so far. DSIT found 75 per cent of UK adopters reported improved productivity but only 12 per cent reported higher revenue, and 77 per cent had seen no revenue change. PwC found 21 per cent of UK chief executives saying AI had increased revenue in the last twelve months.

We have a working prototype. Can you take it to production?

Usually, and the first step is a short review of what production actually requires: integration, cost per user, accuracy thresholds and ownership. Sometimes the prototype is a good foundation and sometimes rebuilding the core is faster. We will tell you which, and our implementation service is £4,000 to £15,000 fixed.

How accurate does an AI feature need to be?

It depends entirely on what happens when it is wrong. A feature that drafts something a human reviews can tolerate far more error than one that acts unattended. Agreeing that threshold in advance, as a number measured on a test set that includes the messy inputs, with a routing rule for anything below it, is what makes shipping possible.

Should we build the pilot ourselves or use a partner?

MIT’s sample found external partnerships with customised tools reached deployment about 67 per cent of the time against about 33 per cent for internal builds, though the figures are self-reported. A reasonable split is to run the tool-level experiments yourself and bring in a partner for the production stage, which is where integration and cost control live.

How much should a pilot cost?

On our ladder an audit that doubles as the pilot is £2,500 to £6,000, credited against the build; UK boutiques generally charge £1,500 to £5,000 for the equivalent, and a first implementation is £4,000 to £15,000. If a proposed pilot costs more than the production build would, ask why, because the pilot is supposed to be the cheap part.

Why does cost stop projects specifically?

Because it scales with usage rather than headcount, so a feature that is cheap in a pilot with ten users can be alarming at a thousand. Modelling cost per user before scaling, and putting a tested ceiling on it, turns that from a surprise into a decision. The model is usually the smallest line; messages and voice minutes are the large ones.

What should we measure before starting an AI pilot?

Baseline hours per item, monthly volume, an accuracy threshold and what happens on an error, a test set that includes messy inputs, the integration path, a run-cost ceiling in pounds, a named owner and a stop rule. If the supplier cannot help you fill those in during scoping, do not start.

Read next

AI consultancy in London →What an AI feature costs to run →What an AI readiness audit should contain →Your data when a supplier builds your software →What an AI consultant costs in the UK →Our prices, in full →

Get a stalled pilot moving

Tell us what you built and where it stopped. We will tell you what production actually requires, what it costs to get there, and what it will cost to run.

Start a conversation →