Guides

Which Claude model for which job: a UK business guide to Haiku, Sonnet, Opus and Fable.

By Zain M · 6 October 2026 · 11 min read

For a UK business the rule is to use the smallest Claude model that passes your own test on your own inputs, and to reserve the larger ones for the steps that need them. At September 2026 list prices Haiku 4.5 is $1 per million input tokens and $5 output, Sonnet 5 is $2 and $10, Opus 5.5 is $4 and $20 and Fable 5.1 is $10 and $50, about £0.75 to £7.55 per million input tokens. Haiku handles reading, sorting, extraction and routine replies; Sonnet handles most drafting and analysis; Opus handles long, careful reasoning and complex code; Fable is for the hardest problems where the cost is justified. From running production systems, the frontier model handles roughly one call in eighty and the small model does the rest, which is why the per-token headline overstates what a working system costs.

DefaultThe smallest model that passes your test
Input price range$1 to $10 per million tokens
In productionFrontier model on roughly one call in eighty

The four models and what they cost

Anthropic’s current lineup on the API is Haiku 4.5, Sonnet 5, Opus 5.5 and Fable 5.1, priced per million tokens with input and output charged separately. Prompt caching is included on all of them and batch processing halves the price for work that can wait. In the chat products the plan decides which models are available and the usage allowance; on the API you choose per request, which is where the routing decisions below apply.

ModelInput, per million tokensOutput, per million tokensAbout in pounds (input / output)Best at
Haiku 4.5$1$5£0.75 / £3.77Reading, sorting, extraction, routine replies, high volume
Sonnet 5$2$10£1.51 / £7.55Most drafting, analysis and code; the workhorse
Opus 5.5$4$20£3.02 / £15.09Long careful reasoning, complex code, hard documents
Fable 5.1$10$50£7.55 / £37.73The hardest problems where the cost is justified

Anthropic list prices, read 25 September 2026, at USD 1 = GBP 0.7545. Batch processing is half price; prompt caching reduces repeated input cost.

What a million tokens actually is

A token is roughly three-quarters of an English word. A typical business email is a few hundred tokens; a CV or an invoice a few thousand; a 50-page contract tens of thousands. A million input tokens is therefore something like 300 CVs or a dozen long contracts, and on Haiku that costs about 75 pence. The output side is dearer per token but a summary or a decision is usually far shorter than the input. This is why document-reading automation is cheap to run and why our running-cost guide says the channel, a text message or a minute of telephony, usually costs more than the model behind it.

Job by job

Reading and sorting at volume, applications, invoices, enquiries, tickets: Haiku, with a test on fifty real examples to prove it meets the threshold, and an escalation to Sonnet for the ones it flags as uncertain. Drafting in the firm’s voice, summarising a document, answering from a knowledge base: Sonnet, which is where most business work should sit. Long, careful analysis of a contract, a board pack or a codebase, or a multi-step agent that must not make mistakes: Opus. Fable for the small number of problems where a wrong answer is expensive enough to justify ten times the price, and only after Opus has been tried. Voice and phone: the fast, cheap model, because latency is the experience and the conversation is short.

JobStart withEscalate toWhy
Screen or extract from documents at volumeHaikuSonnet on uncertain casesSpeed, cost, and a person checks exceptions anyway
Draft, summarise, answer from your documentsSonnetOpus for the hard onesQuality per pound
Analyse a long contract or board packOpusFable if the stakes justify itHolds the thread across a long input
Run an agent across several systemsSonnet for steps, Opus for planningFable rarelyMost steps are simple; the plan is not
Answer the phoneHaikuHand to a personLatency is the experience

Routing: how a working system uses all four

The design that keeps a system cheap is routing. Every request goes first to the smallest model that could handle it; the model or a simple rule decides whether the case is routine or hard; hard cases go up one tier; and a person sees the exceptions. Across the platform we run, that pattern means the large frontier model handles roughly one call in eighty, and moving one high-volume feature from the frontier model to the small one made it several times cheaper and faster with no drop in success. Cost is attributed per request to a feature and a customer so the routing can be tuned from evidence, and spending ceilings are enforced in case it goes wrong.

How to choose for your own case

Do not choose from a benchmark. Take fifty real inputs from the process, with the right answer known for each, and run them through Haiku, then Sonnet, then Opus. Score each on your own acceptance threshold. Pick the smallest model that passes, note the cases it fails, and route those up a tier. The whole exercise costs pence on the API and settles the decision in an afternoon, and it is what we do in the discovery step of every build.

Common questions

Which Claude model should a business use?

The smallest that passes your own test on your own inputs. Haiku for reading and sorting at volume, Sonnet for most drafting and analysis, Opus for long careful reasoning and complex code, Fable only where the stakes justify the price.

How much does the Claude API cost?

Haiku 4.5 is $1 per million input tokens and $5 output; Sonnet 5 $2 and $10; Opus 5.5 $4 and $20; Fable 5.1 $10 and $50, at September 2026 list prices. Batch processing halves it.

What is a token?

Roughly three-quarters of a word. A CV or an invoice is a few thousand tokens; a million tokens is something like 300 CVs, which on Haiku costs about 75 pence.

Is the most expensive model the best choice?

Rarely for business work. Most tasks are handled well by the smaller models, and a well-routed system sends the frontier model roughly one call in eighty. Pay for the large model on the steps that need it.

What is model routing?

Sending each request first to the smallest capable model and escalating hard cases up a tier, with a person seeing the exceptions. It is the single biggest lever on running cost.

Does the chat plan let me choose models?

The plan decides which models are available and the usage allowance. Per-request choice is an API decision, which applies to software you build rather than to chat.

Want the model chosen on your own inputs?

The discovery step runs your real cases through each model against your own threshold, then quotes the build with the running cost stated per feature.

Start a conversation →