How to measure AI return in 90 days in a UK business, with numbers you can defend.
By Zain M · 29 September 2026 · 12 min read
To measure AI return in 90 days, pick one process with real volume, measure its baseline for two weeks before anything is built, define four metrics in advance (hours per unit, time to first response, error or rework rate, and cost per unit including the AI running cost), run the automation for eight weeks with a person still deciding, and compare. UK evidence says this discipline is what separates firms that see a return from those that do not: DSIT found 75 per cent of adopters reported better productivity but only 12 per cent higher revenue, and MIT found most enterprise pilots never produce a measurable return. The number you defend at day 90 is hours returned times a loaded hourly cost, minus build and running cost, on one process, not a company-wide estimate.
Why most AI return is unmeasured, not absent
The UK figures are consistent. DSIT’s AI Adoption Research, a 3,500-business study published in February 2026, found that 75 per cent of adopters reported improved workforce productivity, 12 per cent reported higher revenue and 77 per cent had seen no revenue change. PwC’s survey of UK chief executives in January 2026 found 21 per cent saying AI had increased revenue in the previous twelve months, with 74 per cent reporting little or no change. MIT NANDA’s 2025 study of enterprise deployments found 95 per cent of organisations getting no measurable return from generative AI, and Gartner predicted that at least 30 per cent of generative AI projects would be abandoned after proof of concept, naming unclear business value as one of four causes.
Read together, those figures describe a measurement problem as much as a technology one. Productivity gains are real and widely reported; revenue effects are rare because most SME automation removes cost rather than adding sales; and the projects that get abandoned are the ones where nobody wrote down, before the build, what would count as working. The method below exists to write that down.
The 90-day method
Weeks one and two are the baseline, and they happen before any supplier is chosen. Pick one process with real volume: reading applications, chasing invoices, answering the enquiry line, keying supplier documents. Count it. How many units a week, how many hours each unit takes across everyone who touches it, how long from arrival to first action, how often it has to be redone, and what it costs per unit at a loaded hourly rate. Write the four numbers down and get two people to agree them.
Weeks three to ten are the automation running, with a person still deciding anything that affects a customer or an individual. Measure the same four numbers every week, and add the fifth: the AI running cost per unit, from the provider invoices, which any competent build reports per feature from day one. Week eleven is the comparison and week twelve is the decision: extend, change or stop. A stop is a legitimate outcome and cheaper than a slow fade.
| Metric | Baseline, weeks 1 to 2 | Live, weeks 3 to 10 | What it tells you |
|---|---|---|---|
| Hours per unit | Everyone who touches it, measured | Same measure, weekly | The cost actually removed |
| Time to first action | Arrival to first human response | Arrival to first response, human or system | The customer-side gain |
| Rework or error rate | Units redone or corrected | Same, including AI errors caught | Whether quality held |
| Cost per unit | Hours times loaded rate | Hours times rate, plus AI running cost per unit | The number you defend |
Define all four before the build. A metric added afterwards is an argument, not a measurement.
A worked example
A firm receives 400 supplier invoices a month, each keyed and matched by hand at an average of nine minutes across the people who touch it, at a loaded cost of £28 an hour. Baseline: 60 hours a month, £1,680, with a rework rate of one in twenty and an average of three days from arrival to posting. The automation reads each invoice into the ledger, matches it to the order and flags exceptions for a person. After eight weeks: 14 hours a month of human time on exceptions and review, £392, rework one in fifty, same-day posting, and an AI running cost of about £40 a month at published prices for a small model on 400 documents of a few thousand tokens each. Monthly saving £1,248; annualised about £15,000 against a first implementation at £4,000 to £15,000. Payback inside the first year on one process, with the evidence gathered before anyone was asked to believe it.
The numbers in that example are illustrative. The structure is not: the saving is hours times rate, the cost is build plus running cost, and both sides were measured rather than estimated.
The mistakes that produce a number nobody believes
Estimating the baseline from memory instead of measuring it, which produces a saving that finance rejects in the first meeting. Counting the time of the person who used to do the task without asking what they do now, which is fine if it is billable work and an argument if it is not. Leaving the running cost out, which is how a saving turns into a bill. Measuring a company-wide effect from a single-process pilot. And declaring the pilot a success at week four because the demo was impressive, which is the road to the abandoned projects in the surveys above.
What to do at day 90
Extend if the four metrics held and the cost per unit fell by enough to pay for the build within a year. Change if the quality metric moved the wrong way, which usually means the exceptions path needs work rather than the model. Stop if the volume was never there, and write down why, so the next idea is measured against it. Then pick the second process, and run the same twelve weeks. A firm that does this four times in a year has a portfolio of measured returns rather than an AI strategy, and it is a better position to be in.
Common questions
How do you measure ROI on AI?
Measure one process before and after: hours per unit, time to first action, rework rate and cost per unit including the AI running cost. The return is hours returned times a loaded rate, minus build and running cost, on that process, over a defined window.
How long does it take to see a return from AI?
On a single high-volume process, the evidence is visible within eight weeks of going live, and payback is often inside the first year. Company-wide returns take longer and are harder to attribute, which is why the surveys show productivity gains without revenue change.
What is a good ROI for an AI project?
Payback within twelve months on a first process is a sensible bar for a UK SME. The published surveys show most adopters get productivity gains and few get revenue gains, so measure cost removed rather than sales added unless the process is sales.
Why do most AI pilots fail to show a return?
Because nobody wrote down the baseline or the success measure before the build. Gartner lists unclear business value as a cause of abandonment; MIT found most enterprise pilots never reach measurable return. A measured baseline fixes most of it.
Should the AI running cost be included in ROI?
Yes, per unit, from the provider invoices. A saving that leaves out the running cost is not a saving. Any competent build reports cost per feature from day one.
What should a UK SME measure first?
The process with the most volume and the most retyping, chasing or reading: invoices, applications, enquiries, documents. It is where the hours are and where the baseline is easiest to count.
Want the baseline measured for you?
The AI adoption audit measures the processes worth automating, ranks them by return with the running cost included, and is credited in full against the first build.