Which Claude model for which job: a UK business guide to Haiku, Sonnet, Opus and Fable.
By Zain M · 6 October 2026 · 11 min read
For a UK business the rule is to use the smallest Claude model that passes your own test on your own inputs, and to reserve the larger ones for the steps that need them. At September 2026 list prices Haiku 4.5 is $1 per million input tokens and $5 output, Sonnet 5 is $2 and $10, Opus 5.5 is $4 and $20 and Fable 5.1 is $10 and $50, about £0.75 to £7.55 per million input tokens. Haiku handles reading, sorting, extraction and routine replies; Sonnet handles most drafting and analysis; Opus handles long, careful reasoning and complex code; Fable is for the hardest problems where the cost is justified. From running production systems, the frontier model handles roughly one call in eighty and the small model does the rest, which is why the per-token headline overstates what a working system costs.
The four models and what they cost
Anthropic’s current lineup on the API is Haiku 4.5, Sonnet 5, Opus 5.5 and Fable 5.1, priced per million tokens with input and output charged separately. Prompt caching is included on all of them and batch processing halves the price for work that can wait. In the chat products the plan decides which models are available and the usage allowance; on the API you choose per request, which is where the routing decisions below apply.
| Model | Input, per million tokens | Output, per million tokens | About in pounds (input / output) | Best at |
|---|---|---|---|---|
| Haiku 4.5 | $1 | $5 | £0.75 / £3.77 | Reading, sorting, extraction, routine replies, high volume |
| Sonnet 5 | $2 | $10 | £1.51 / £7.55 | Most drafting, analysis and code; the workhorse |
| Opus 5.5 | $4 | $20 | £3.02 / £15.09 | Long careful reasoning, complex code, hard documents |
| Fable 5.1 | $10 | $50 | £7.55 / £37.73 | The hardest problems where the cost is justified |
Anthropic list prices, read 25 September 2026, at USD 1 = GBP 0.7545. Batch processing is half price; prompt caching reduces repeated input cost.
What a million tokens actually is
A token is roughly three-quarters of an English word. A typical business email is a few hundred tokens; a CV or an invoice a few thousand; a 50-page contract tens of thousands. A million input tokens is therefore something like 300 CVs or a dozen long contracts, and on Haiku that costs about 75 pence. The output side is dearer per token but a summary or a decision is usually far shorter than the input. This is why document-reading automation is cheap to run and why our running-cost guide says the channel, a text message or a minute of telephony, usually costs more than the model behind it.
Job by job
Reading and sorting at volume, applications, invoices, enquiries, tickets: Haiku, with a test on fifty real examples to prove it meets the threshold, and an escalation to Sonnet for the ones it flags as uncertain. Drafting in the firm’s voice, summarising a document, answering from a knowledge base: Sonnet, which is where most business work should sit. Long, careful analysis of a contract, a board pack or a codebase, or a multi-step agent that must not make mistakes: Opus. Fable for the small number of problems where a wrong answer is expensive enough to justify ten times the price, and only after Opus has been tried. Voice and phone: the fast, cheap model, because latency is the experience and the conversation is short.
| Job | Start with | Escalate to | Why |
|---|---|---|---|
| Screen or extract from documents at volume | Haiku | Sonnet on uncertain cases | Speed, cost, and a person checks exceptions anyway |
| Draft, summarise, answer from your documents | Sonnet | Opus for the hard ones | Quality per pound |
| Analyse a long contract or board pack | Opus | Fable if the stakes justify it | Holds the thread across a long input |
| Run an agent across several systems | Sonnet for steps, Opus for planning | Fable rarely | Most steps are simple; the plan is not |
| Answer the phone | Haiku | Hand to a person | Latency is the experience |
Routing: how a working system uses all four
The design that keeps a system cheap is routing. Every request goes first to the smallest model that could handle it; the model or a simple rule decides whether the case is routine or hard; hard cases go up one tier; and a person sees the exceptions. Across the platform we run, that pattern means the large frontier model handles roughly one call in eighty, and moving one high-volume feature from the frontier model to the small one made it several times cheaper and faster with no drop in success. Cost is attributed per request to a feature and a customer so the routing can be tuned from evidence, and spending ceilings are enforced in case it goes wrong.
How to choose for your own case
Do not choose from a benchmark. Take fifty real inputs from the process, with the right answer known for each, and run them through Haiku, then Sonnet, then Opus. Score each on your own acceptance threshold. Pick the smallest model that passes, note the cases it fails, and route those up a tier. The whole exercise costs pence on the API and settles the decision in an afternoon, and it is what we do in the discovery step of every build.
Common questions
Which Claude model should a business use?
The smallest that passes your own test on your own inputs. Haiku for reading and sorting at volume, Sonnet for most drafting and analysis, Opus for long careful reasoning and complex code, Fable only where the stakes justify the price.
How much does the Claude API cost?
Haiku 4.5 is $1 per million input tokens and $5 output; Sonnet 5 $2 and $10; Opus 5.5 $4 and $20; Fable 5.1 $10 and $50, at September 2026 list prices. Batch processing halves it.
What is a token?
Roughly three-quarters of a word. A CV or an invoice is a few thousand tokens; a million tokens is something like 300 CVs, which on Haiku costs about 75 pence.
Is the most expensive model the best choice?
Rarely for business work. Most tasks are handled well by the smaller models, and a well-routed system sends the frontier model roughly one call in eighty. Pay for the large model on the steps that need it.
What is model routing?
Sending each request first to the smallest capable model and escalating hard cases up a tier, with a person seeing the exceptions. It is the single biggest lever on running cost.
Does the chat plan let me choose models?
The plan decides which models are available and the usage allowance. Per-request choice is an API decision, which applies to software you build rather than to chat.
Want the model chosen on your own inputs?
The discovery step runs your real cases through each model against your own threshold, then quotes the build with the running cost stated per feature.