What an AI feature actually costs to run, per user, per month.
By Zain M · Updated 14 September 2026 · 13 min read
AI running costs scale with usage, not headcount, so the meaningful figure is cost per active user per month. At September 2026 prices a small model costs $0.05 to $0.25 per million input tokens and a frontier model $5 to $10, while one UK text message costs $0.056. In our products the model is a single-digit share of the bill.
Why the monthly total tells you nothing, and what paying our own bills taught us
A database does not get more expensive because one customer had a busy Tuesday. A language model does. That single difference is why AI budgeting defeats teams who are otherwise good at forecasting: every other line in the software budget is a licence or a server, fixed in advance, and this one is a meter that runs faster when the product is used more, which is exactly when nobody wants to be told to slow down.
The number that matters is what one active user costs you in a month, because that is the figure you can multiply. A monthly total tells you what happened. A per user figure tells you what happens if you double, and it is the figure that lets a finance director compare the AI feature with the subscription price it sits inside, which is the comparison that decides whether the feature survives.
This revision adds what the original lacked: the published prices, read on the providers’ pages on 14 September 2026, and a worked example at stated volumes so that a buyer can do the sum before a supplier does it for them. It also adds the channel prices, because the sum is wrong without them. The first-party findings from our own products are unchanged from the original, because they have not changed.
We run two AI products in production, a recruitment platform and a consumer careers product, and attribute every model call and every message to a feature at the moment it is made. That ledger is the reason this guide can say what it says: the findings below are measured, not estimated. Three things in that record changed how we advise clients, and none of them is what the model providers’ marketing would lead you to expect.
What the models cost today, per million tokens
These are the list prices on the three providers’ pricing pages on 14 September 2026, in US dollars, for standard (non-batch) requests. Two things to notice. The spread between the cheapest and the most expensive model on input tokens is more than a hundredfold. And every provider now sells a small, a medium and a large model, so "which model" is a per-feature decision with a per-feature price, which is what routing means.
| Provider and model | Input, per 1M tokens | Output, per 1M tokens | Note |
|---|---|---|---|
| OpenAI gpt-5-nano | $0.05 | $0.40 | Cached input $0.005 |
| OpenAI gpt-5-mini | $0.25 | $2.00 | Cached input $0.025 |
| OpenAI gpt-5 | $1.25 | $10.00 | Cached input $0.125 |
| OpenAI gpt-4.1 / mini / nano | $2.00 / $0.40 / $0.10 | $8.00 / $1.60 / $0.40 | Previous generation, still listed |
| OpenAI o3 / o4-mini | $2.00 / $1.10 | $8.00 / $4.40 | Reasoning models |
| OpenAI gpt-realtime-2.1 / mini (audio) | $32.00 / $10.00 | $64.00 / $20.00 | Audio tokens; gpt-live-1 is priced at $0.05 a minute |
| Anthropic Haiku 4.5 | $1.00 | $5.00 | Cache read $0.10 |
| Anthropic Sonnet 5 | $2.00 | $10.00 | Cache read $0.20 |
| Anthropic Opus 5 | $5.00 | $25.00 | Cache read $0.50 |
| Anthropic Fable 5.1 | $10.00 | $50.00 | Cache read $0.25; batch 50% off across models |
| Google Gemini 2.5 Flash-Lite | $0.10 | $0.40 | Caching $0.01 |
| Google Gemini 2.5 Flash / 3.5 Flash-Lite | $0.30 | $2.50 | Same price for both |
| Google Gemini 3.8 Flash | $0.75 | $3.75 | Rises to $1.50 and $7.50 from 1 Jan 2027 |
| Google Gemini 3.5 Flash | $1.50 | $9.00 | |
| Google Gemini 2.5 Pro / 3.1 Pro preview | $1.25 / $2.00 | $10.00 / $12.00 | Prompts over 200k tokens cost more |
Read on the OpenAI, Anthropic and Google pricing pages on 14 September 2026. US dollars, excluding VAT. Prices change; the date is the thing to check.
What the channels cost: SMS and voice in the UK
A feature that sends a message or makes a call spends more on the sending than on the thinking, and these are the published UK rates from the platform most agents are built on. The figures explain the first of our three lessons on their own: one outbound text costs more than a thousand input tokens on every small model in the table above, and a minute of a call to a mobile costs more than ten thousand.
| Twilio, United Kingdom | Published price | Note |
|---|---|---|
| Outbound SMS | $0.056 per message | To mobiles and from alphanumeric sender IDs; carrier fees may apply |
| Inbound SMS | $0.0075 per message | |
| Alphanumeric sender ID | Free | A mobile number is $2.50 a month, a local number $1.15 |
| Outbound voice to UK mobiles | $0.0305 per minute | Landlines $0.0158 |
| Inbound voice | $0.0100 per minute | Plus $3.50 a month for a local number |
Read on Twilio’s UK SMS and voice pricing pages on 14 September 2026, pay-as-you-go, US dollars. Voice AI adds the model’s audio cost on top of the carrier minute.
A worked example at stated volumes
A hypothetical, using only the prices above. Suppose a business application has 1,000 active users, each making 30 requests a month, and each request sends about 2,000 tokens of input and receives 300 of output. That is 30,000 requests, 60 million input tokens and 9 million output tokens a month. Two thousand input tokens is roughly a page and a half of text, so this is a feature that reads a document or a record and returns a short answer; a feature that writes long prose would have the ratio the other way round.
On gpt-5-nano the month costs $3.00 plus $3.60, or $6.60: under a cent per user. On gpt-5-mini it is $15 plus $18, or $33. On Haiku 4.5, $105. On Sonnet 5, $120 plus $90, or $210, about 21 cents per user. On Opus 5, $525. On Fable 5.1, $600 plus $450, or $1,050, just over a dollar per user. The same feature, the same users, a 160-fold range depending on one decision. Now apply our own routing ratio as an illustration: if one request in eighty needs the frontier model and the rest run on gpt-5-mini, the month is about $46, and the per-user figure is under five cents.
Then add a channel. If the feature sends each user one text message a month, that is 1,000 messages at $0.056, or $56: more than the entire model bill on the small model. Five texts a month, $280, exceeds Sonnet 5. If it makes each user one three-minute call to a mobile, the carrier minutes are 3,000 at $0.0305, or $91.50, before the voice model, which at gpt-live-1’s $0.05 a minute adds $150. The channel, not the model, is where the per-user figure is decided, which is why we model it first and cap it separately. Convert at the rate on the day and add VAT; the ordering does not change.
Where the pricing pages hide the cost
The headline per-million figures are the start of the sum, not the end of it, and each provider’s page carries four qualifications that change the answer. Batch processing: Anthropic advertises 50 per cent off for batch requests across its models, which suits anything that does not need an answer within seconds, such as overnight scoring of a day’s documents. Cached input: OpenAI lists cached input on gpt-5 at $0.125 against $1.25, Anthropic cache reads on Sonnet 5 at $0.20 against $2, Google caching on Gemini 2.5 Flash at $0.03 against $0.30, and all of them charge for it only if the same context is actually resent.
Long context: Google prices prompts over 200,000 tokens higher, with Gemini 2.5 Pro rising from $1.25 to $2.50 on input and from $10 to $15 on output, so a feature that stuffs a whole archive into every request pays twice. Data residency: OpenAI adds 10 per cent for eligible models released on or after 5 March 2026 when data is pinned to a region such as the United Kingdom, and Anthropic prices US-only inference at 1.1 times the standard rate, so the compliance decision in our data guide has a line in this one too.
And prices move in both directions. Google’s page states that Gemini 3.8 Flash rises from $0.75 and $3.75 to $1.50 and $7.50 per million on 1 January 2027, a doubling announced in advance; other models on the same page are cheaper than their predecessors were. A budget built on today’s prices needs a date on it, and attribution is what lets you re-run the sum when the date passes.
The three controls that make it predictable
Ceilings and instrumentation, in practice
Attribution means every model call and every message is written to a ledger with the customer, the feature, the model and the token or unit count before the response is returned, so the cost of any feature for any customer in any month is a query rather than an estimate. The provider dashboards will not do this for you; they know the organisation, not the customer. Cached-input pricing, which every provider now publishes at a fraction of the standard rate, only helps if you can see which features resend the same context.
Ceilings are set per user per day and per month, in the currency of the thing being capped: tokens for the model, messages for SMS, minutes for voice. A cap that is set but never triggered is a hope. We fire ours deliberately in production, with a test account, and watch the feature degrade gracefully into a queue or a fallback rather than an error, because a customer who hits a ceiling on a Friday evening should not find a broken screen.
Retries belong inside the cap. Anything that runs unattended and retries on failure can multiply cost without changing behaviour, so the expensive work sits inside a retry boundary that counts against the ceiling, not outside it. And long inputs cost more than long outputs in most pricing models, so sending an entire document when a section would do is the most common avoidable expense.
What customers will wait for
Speed is a cost decision too, because the fast model is usually the cheap one. From our own systems: a CV is read and scored in under two seconds in the typical case, a full CV import takes around twenty, and the slowest features, which reason over a whole profile, take the better part of a minute. The two-second feature runs on the small model; the minute-long one is where the frontier model earns its price, and the user is warned.
The rule we apply is that anything over ten seconds needs visible progress or the user assumes it has broken, and anything over thirty needs to run in the background with a notification when it finishes. Design the wait before you pick the model, not after, because a feature that is moved to a slower model to save money without redesigning the wait will lose users faster than it saves pounds.
The three lessons, and what good looks like
First, the model is not the bill: budget the channels around it, because a text at $0.056 costs more than a thousand tokens on any small model. Second, the small model does most of the work: decide the model per feature, and expect the frontier model to be the exception. Third, routing beats prompt tuning: a swap that made one feature five times cheaper and four times faster is the largest single saving we have made, and it took an afternoon.
What good looks like is this. You can state the cost per active user per month from a query rather than an estimate. You know which feature is the most expensive and why. Your spending ceiling has been triggered deliberately at least once so you know it works. And the usage that drives your invoicing and the usage that drives your costs come from the same record, so they cannot drift apart.
If you are buying rather than building, the same four sentences are the questions to put to the supplier, in that order. A supplier who can answer the first from a query, the second by feature, the third with a date and the fourth with a diagram has built cost control in. One who answers with a monthly estimate has built a feature and left the bill to you.
Method and sources
Model and channel prices were read on the OpenAI, Anthropic, Google and Twilio pricing pages on 14 September 2026 and are quoted in the currency each vendor bills in, excluding VAT; they change, and the date is the thing to check. The worked example uses hypothetical volumes stated in the text and the published prices only. The first-party findings, the model’s single-digit share of variable AI spend, the roughly one-in-eighty frontier ratio, the five-times-cheaper and four-times-faster routing swap, and the response times, are from our own production ledger and are published here as shares and ratios; nothing on this page states a volume, a total or a cost per feature.
Common questions
What is a normal cost per user for an AI feature?
It varies enormously by feature. At September 2026 prices, 30 text requests a month of 2,000 input and 300 output tokens cost under a cent per user on a small model and about a dollar on the largest, before any messaging. Anything involving voice, long documents or heavy automated processing can be an order of magnitude higher, which is why per-feature attribution matters more than a benchmark.
How much do AI models cost per million tokens in 2026?
On the providers’ pages on 14 September 2026: OpenAI gpt-5-nano $0.05 input and $0.40 output, gpt-5-mini $0.25 and $2, gpt-5 $1.25 and $10; Anthropic Haiku 4.5 $1 and $5, Sonnet 5 $2 and $10, Opus 5 $5 and $25, Fable 5.1 $10 and $50; Google Gemini 2.5 Flash-Lite $0.10 and $0.40, 3.8 Flash $0.75 and $3.75, 2.5 Pro $1.25 and $10.
How much does an SMS cost in the UK through an API?
Twilio’s published UK rate on 14 September 2026 is $0.056 per outbound message and $0.0075 inbound, with alphanumeric sender IDs free and a mobile number at $2.50 a month. One message costs more than a thousand tokens on any small model, which is why messaging is usually the largest line in an agent that contacts people.
Can we cap what a customer costs us?
Yes, per user daily and monthly ceilings are straightforward to enforce and are the difference between a predictable line and an open-ended one. Set them in the unit being spent, tokens, messages or minutes, decide what the user sees when the cap is hit, and test that the cap actually stops work rather than assuming it would. A limit nobody has seen fire is an assumption.
Does a cheaper model always mean worse results?
No. For classification, extraction and routine drafting, smaller models often perform indistinguishably at a fraction of the cost. When we moved one of our own high volume features to the small model it got roughly five times cheaper and four times faster with no drop in success rate. The skill is knowing which work genuinely needs the expensive model, which is what routing decides.
What does prompt caching save?
Every provider now prices cached input at a fraction of the standard rate: OpenAI lists cached input on gpt-5 at $0.125 against $1.25, Anthropic cache reads on Sonnet 5 at $0.20 against $2, Google caching on 2.5 Flash at $0.03 against $0.30. It only helps for context that repeats, which attribution will show you.
Should we ask a supplier which model they use?
Ask which models they use and why, per feature, and what the fallback is. A supplier who defaults everything to the largest model has not done the routing work and you will pay for that every month, at up to 200 times the price of the smallest model on input tokens.
How much does AI voice cost per minute?
The carrier minute to a UK mobile is $0.0305 on Twilio’s published rate, and the model’s audio is on top: OpenAI lists gpt-live-1 at $0.05 a minute and its realtime audio models at $10 to $32 per million input tokens. Budget voice separately from text and cap it per user, and remember that in our production data most outbound calls reached voicemail.
What is the cheapest way to run an AI feature?
Route by feature so routine work runs on a small model, cache repeated context, keep inputs short, use batch pricing for anything that can wait, and put the expensive work inside a retry boundary that counts against a ceiling. In our experience routing alone changed the bill more than every prompt optimisation combined, by roughly five times on the feature we moved.
Find out what yours costs
If you are running AI features and cannot answer the per user question, that is the audit. We instrument it, show you the number, and put a ceiling on it.