Guides

What an AI feature actually costs to run, per user, per month.

By Zain M · Updated 14 September 2026 · 13 min read

AI running costs scale with usage, not headcount, so the meaningful figure is cost per active user per month. At September 2026 prices a small model costs $0.05 to $0.25 per million input tokens and a frontier model $5 to $10, while one UK text message costs $0.056. In our products the model is a single-digit share of the bill.

MeasurePer active user, per month
Small to frontier model spreadUp to 200 times on input tokens
Biggest leverModel routing
Most skippedTesting the ceiling

Why the monthly total tells you nothing, and what paying our own bills taught us

A database does not get more expensive because one customer had a busy Tuesday. A language model does. That single difference is why AI budgeting defeats teams who are otherwise good at forecasting: every other line in the software budget is a licence or a server, fixed in advance, and this one is a meter that runs faster when the product is used more, which is exactly when nobody wants to be told to slow down.

The number that matters is what one active user costs you in a month, because that is the figure you can multiply. A monthly total tells you what happened. A per user figure tells you what happens if you double, and it is the figure that lets a finance director compare the AI feature with the subscription price it sits inside, which is the comparison that decides whether the feature survives.

This revision adds what the original lacked: the published prices, read on the providers’ pages on 14 September 2026, and a worked example at stated volumes so that a buyer can do the sum before a supplier does it for them. It also adds the channel prices, because the sum is wrong without them. The first-party findings from our own products are unchanged from the original, because they have not changed.

We run two AI products in production, a recruitment platform and a consumer careers product, and attribute every model call and every message to a feature at the moment it is made. That ledger is the reason this guide can say what it says: the findings below are measured, not estimated. Three things in that record changed how we advise clients, and none of them is what the model providers’ marketing would lead you to expect.

01The model is not the bill. On the recruitment platform, language model calls were a single digit share of the variable spend on AI features. Text messages to candidates were the largest line by a wide margin, with telephony second. Teams who budget for the model and forget the channels around it are surprised every month.
02The small model does most of the work. Across that platform, the large frontier model handled roughly one call in eighty. Everything else, including reading and scoring every CV, ran on the small model, because classification and extraction do not need the expensive one. On the consumer product the mix is the other way round, because the output is prose a person reads and judges. The right mix is decided per feature, never per company.
03Routing is worth more than prompt tuning. We moved one high volume feature from the frontier model to the small one. It became roughly five times cheaper and four times faster with no change in success rate. Nothing else we did to that feature came close.

What the models cost today, per million tokens

These are the list prices on the three providers’ pricing pages on 14 September 2026, in US dollars, for standard (non-batch) requests. Two things to notice. The spread between the cheapest and the most expensive model on input tokens is more than a hundredfold. And every provider now sells a small, a medium and a large model, so "which model" is a per-feature decision with a per-feature price, which is what routing means.

Provider and modelInput, per 1M tokensOutput, per 1M tokensNote
OpenAI gpt-5-nano$0.05$0.40Cached input $0.005
OpenAI gpt-5-mini$0.25$2.00Cached input $0.025
OpenAI gpt-5$1.25$10.00Cached input $0.125
OpenAI gpt-4.1 / mini / nano$2.00 / $0.40 / $0.10$8.00 / $1.60 / $0.40Previous generation, still listed
OpenAI o3 / o4-mini$2.00 / $1.10$8.00 / $4.40Reasoning models
OpenAI gpt-realtime-2.1 / mini (audio)$32.00 / $10.00$64.00 / $20.00Audio tokens; gpt-live-1 is priced at $0.05 a minute
Anthropic Haiku 4.5$1.00$5.00Cache read $0.10
Anthropic Sonnet 5$2.00$10.00Cache read $0.20
Anthropic Opus 5$5.00$25.00Cache read $0.50
Anthropic Fable 5.1$10.00$50.00Cache read $0.25; batch 50% off across models
Google Gemini 2.5 Flash-Lite$0.10$0.40Caching $0.01
Google Gemini 2.5 Flash / 3.5 Flash-Lite$0.30$2.50Same price for both
Google Gemini 3.8 Flash$0.75$3.75Rises to $1.50 and $7.50 from 1 Jan 2027
Google Gemini 3.5 Flash$1.50$9.00
Google Gemini 2.5 Pro / 3.1 Pro preview$1.25 / $2.00$10.00 / $12.00Prompts over 200k tokens cost more

Read on the OpenAI, Anthropic and Google pricing pages on 14 September 2026. US dollars, excluding VAT. Prices change; the date is the thing to check.

What the channels cost: SMS and voice in the UK

A feature that sends a message or makes a call spends more on the sending than on the thinking, and these are the published UK rates from the platform most agents are built on. The figures explain the first of our three lessons on their own: one outbound text costs more than a thousand input tokens on every small model in the table above, and a minute of a call to a mobile costs more than ten thousand.

Twilio, United KingdomPublished priceNote
Outbound SMS$0.056 per messageTo mobiles and from alphanumeric sender IDs; carrier fees may apply
Inbound SMS$0.0075 per message
Alphanumeric sender IDFreeA mobile number is $2.50 a month, a local number $1.15
Outbound voice to UK mobiles$0.0305 per minuteLandlines $0.0158
Inbound voice$0.0100 per minutePlus $3.50 a month for a local number

Read on Twilio’s UK SMS and voice pricing pages on 14 September 2026, pay-as-you-go, US dollars. Voice AI adds the model’s audio cost on top of the carrier minute.

A worked example at stated volumes

A hypothetical, using only the prices above. Suppose a business application has 1,000 active users, each making 30 requests a month, and each request sends about 2,000 tokens of input and receives 300 of output. That is 30,000 requests, 60 million input tokens and 9 million output tokens a month. Two thousand input tokens is roughly a page and a half of text, so this is a feature that reads a document or a record and returns a short answer; a feature that writes long prose would have the ratio the other way round.

On gpt-5-nano the month costs $3.00 plus $3.60, or $6.60: under a cent per user. On gpt-5-mini it is $15 plus $18, or $33. On Haiku 4.5, $105. On Sonnet 5, $120 plus $90, or $210, about 21 cents per user. On Opus 5, $525. On Fable 5.1, $600 plus $450, or $1,050, just over a dollar per user. The same feature, the same users, a 160-fold range depending on one decision. Now apply our own routing ratio as an illustration: if one request in eighty needs the frontier model and the rest run on gpt-5-mini, the month is about $46, and the per-user figure is under five cents.

Then add a channel. If the feature sends each user one text message a month, that is 1,000 messages at $0.056, or $56: more than the entire model bill on the small model. Five texts a month, $280, exceeds Sonnet 5. If it makes each user one three-minute call to a mobile, the carrier minutes are 3,000 at $0.0305, or $91.50, before the voice model, which at gpt-live-1’s $0.05 a minute adds $150. The channel, not the model, is where the per-user figure is decided, which is why we model it first and cap it separately. Convert at the rate on the day and add VAT; the ordering does not change.

Where the pricing pages hide the cost

The headline per-million figures are the start of the sum, not the end of it, and each provider’s page carries four qualifications that change the answer. Batch processing: Anthropic advertises 50 per cent off for batch requests across its models, which suits anything that does not need an answer within seconds, such as overnight scoring of a day’s documents. Cached input: OpenAI lists cached input on gpt-5 at $0.125 against $1.25, Anthropic cache reads on Sonnet 5 at $0.20 against $2, Google caching on Gemini 2.5 Flash at $0.03 against $0.30, and all of them charge for it only if the same context is actually resent.

Long context: Google prices prompts over 200,000 tokens higher, with Gemini 2.5 Pro rising from $1.25 to $2.50 on input and from $10 to $15 on output, so a feature that stuffs a whole archive into every request pays twice. Data residency: OpenAI adds 10 per cent for eligible models released on or after 5 March 2026 when data is pinned to a region such as the United Kingdom, and Anthropic prices US-only inference at 1.1 times the standard rate, so the compliance decision in our data guide has a line in this one too.

And prices move in both directions. Google’s page states that Gemini 3.8 Flash rises from $0.75 and $3.75 to $1.50 and $7.50 per million on 1 January 2027, a doubling announced in advance; other models on the same page are cheaper than their predecessors were. A budget built on today’s prices needs a date on it, and attribution is what lets you re-run the sum when the date passes.

The three controls that make it predictable

FRONTIER MODEL→ SMALL MODELCost per callabout a fifthTime to answerabout a quarterSuccess rateunchangedOne feature, before and after. Nothing else about it changed.
Moving one high volume feature to the small model
01Attribution at the moment of the request, recording which customer, which feature and which model. Retrofitting this later is close to impossible because the information is gone.
02Routing, so routine work runs on economical models and expensive ones are reserved for work that genuinely needs them. This usually changes the bill more than any prompt optimisation, and the swap above is what that looks like in practice.
03A hard ceiling that has been watched stopping something in production. A limit nobody has seen fire is an assumption, not a control.

Ceilings and instrumentation, in practice

Attribution means every model call and every message is written to a ledger with the customer, the feature, the model and the token or unit count before the response is returned, so the cost of any feature for any customer in any month is a query rather than an estimate. The provider dashboards will not do this for you; they know the organisation, not the customer. Cached-input pricing, which every provider now publishes at a fraction of the standard rate, only helps if you can see which features resend the same context.

Ceilings are set per user per day and per month, in the currency of the thing being capped: tokens for the model, messages for SMS, minutes for voice. A cap that is set but never triggered is a hope. We fire ours deliberately in production, with a test account, and watch the feature degrade gracefully into a queue or a fallback rather than an error, because a customer who hits a ceiling on a Friday evening should not find a broken screen.

Retries belong inside the cap. Anything that runs unattended and retries on failure can multiply cost without changing behaviour, so the expensive work sits inside a retry boundary that counts against the ceiling, not outside it. And long inputs cost more than long outputs in most pricing models, so sending an entire document when a section would do is the most common avoidable expense.

What customers will wait for

Speed is a cost decision too, because the fast model is usually the cheap one. From our own systems: a CV is read and scored in under two seconds in the typical case, a full CV import takes around twenty, and the slowest features, which reason over a whole profile, take the better part of a minute. The two-second feature runs on the small model; the minute-long one is where the frontier model earns its price, and the user is warned.

The rule we apply is that anything over ten seconds needs visible progress or the user assumes it has broken, and anything over thirty needs to run in the background with a notification when it finishes. Design the wait before you pick the model, not after, because a feature that is moved to a slower model to save money without redesigning the wait will lose users faster than it saves pounds.

The three lessons, and what good looks like

First, the model is not the bill: budget the channels around it, because a text at $0.056 costs more than a thousand tokens on any small model. Second, the small model does most of the work: decide the model per feature, and expect the frontier model to be the exception. Third, routing beats prompt tuning: a swap that made one feature five times cheaper and four times faster is the largest single saving we have made, and it took an afternoon.

What good looks like is this. You can state the cost per active user per month from a query rather than an estimate. You know which feature is the most expensive and why. Your spending ceiling has been triggered deliberately at least once so you know it works. And the usage that drives your invoicing and the usage that drives your costs come from the same record, so they cannot drift apart.

If you are buying rather than building, the same four sentences are the questions to put to the supplier, in that order. A supplier who can answer the first from a query, the second by feature, the third with a date and the fourth with a diagram has built cost control in. One who answers with a monthly estimate has built a feature and left the bill to you.

Method and sources

Model and channel prices were read on the OpenAI, Anthropic, Google and Twilio pricing pages on 14 September 2026 and are quoted in the currency each vendor bills in, excluding VAT; they change, and the date is the thing to check. The worked example uses hypothetical volumes stated in the text and the published prices only. The first-party findings, the model’s single-digit share of variable AI spend, the roughly one-in-eighty frontier ratio, the five-times-cheaper and four-times-faster routing swap, and the response times, are from our own production ledger and are published here as shares and ratios; nothing on this page states a volume, a total or a cost per feature.

Common questions

What is a normal cost per user for an AI feature?

It varies enormously by feature. At September 2026 prices, 30 text requests a month of 2,000 input and 300 output tokens cost under a cent per user on a small model and about a dollar on the largest, before any messaging. Anything involving voice, long documents or heavy automated processing can be an order of magnitude higher, which is why per-feature attribution matters more than a benchmark.

How much do AI models cost per million tokens in 2026?

On the providers’ pages on 14 September 2026: OpenAI gpt-5-nano $0.05 input and $0.40 output, gpt-5-mini $0.25 and $2, gpt-5 $1.25 and $10; Anthropic Haiku 4.5 $1 and $5, Sonnet 5 $2 and $10, Opus 5 $5 and $25, Fable 5.1 $10 and $50; Google Gemini 2.5 Flash-Lite $0.10 and $0.40, 3.8 Flash $0.75 and $3.75, 2.5 Pro $1.25 and $10.

How much does an SMS cost in the UK through an API?

Twilio’s published UK rate on 14 September 2026 is $0.056 per outbound message and $0.0075 inbound, with alphanumeric sender IDs free and a mobile number at $2.50 a month. One message costs more than a thousand tokens on any small model, which is why messaging is usually the largest line in an agent that contacts people.

Can we cap what a customer costs us?

Yes, per user daily and monthly ceilings are straightforward to enforce and are the difference between a predictable line and an open-ended one. Set them in the unit being spent, tokens, messages or minutes, decide what the user sees when the cap is hit, and test that the cap actually stops work rather than assuming it would. A limit nobody has seen fire is an assumption.

Does a cheaper model always mean worse results?

No. For classification, extraction and routine drafting, smaller models often perform indistinguishably at a fraction of the cost. When we moved one of our own high volume features to the small model it got roughly five times cheaper and four times faster with no drop in success rate. The skill is knowing which work genuinely needs the expensive model, which is what routing decides.

What does prompt caching save?

Every provider now prices cached input at a fraction of the standard rate: OpenAI lists cached input on gpt-5 at $0.125 against $1.25, Anthropic cache reads on Sonnet 5 at $0.20 against $2, Google caching on 2.5 Flash at $0.03 against $0.30. It only helps for context that repeats, which attribution will show you.

Should we ask a supplier which model they use?

Ask which models they use and why, per feature, and what the fallback is. A supplier who defaults everything to the largest model has not done the routing work and you will pay for that every month, at up to 200 times the price of the smallest model on input tokens.

How much does AI voice cost per minute?

The carrier minute to a UK mobile is $0.0305 on Twilio’s published rate, and the model’s audio is on top: OpenAI lists gpt-live-1 at $0.05 a minute and its realtime audio models at $10 to $32 per million input tokens. Budget voice separately from text and cap it per user, and remember that in our production data most outbound calls reached voicemail.

What is the cheapest way to run an AI feature?

Route by feature so routine work runs on a small model, cache repeated context, keep inputs short, use batch pricing for anything that can wait, and put the expensive work inside a retry boundary that counts against a ceiling. In our experience routing alone changed the bill more than every prompt optimisation combined, by roughly five times on the feature we moved.

Read next

AI automation agency in London →What AI automation costs in the UK →Does AI phone calling actually work? →Why most AI pilots never reach production →Our prices, in full →

Find out what yours costs

If you are running AI features and cannot answer the per user question, that is the audit. We instrument it, show you the number, and put a ceiling on it.

Start a conversation →