Svennis AI
10 min read

Opus, Sonnet or Haiku: which Claude model when, matched to each business task

A decision guide to the current Claude models. Match Haiku, Sonnet, Opus and Fable to the job each step of your workflow does, then test on twenty real examples before you commit.

Abstract layered shapes of three sizes moving along parallel paths, suggesting tasks sorted by weight

Opus, Sonnet or Haiku: match the Claude model to the job each step does

The answer to "Opus, Sonnet or Haiku: which Claude model when" is to choose by the job, not by the newest or biggest name. Use Haiku for high-volume sorting and routing. Use Sonnet for most everyday drafting and analysis. Keep Opus for the few hard, high-stakes judgements where a wrong answer costs real money.

Model selection is the choice of which Claude model handles a given step of a workflow. A single business process often has several steps, and each step can use a different model. Reading an email and deciding which queue it belongs to is a different job from writing a careful reply to an angry client.

This guide sits above the four model posts already on this site. It gives you the decision, a table of real small-business jobs and a way to test your choice. For the detail on each model, the individual posts go deeper.

Capabilities, speed and cost: the three criteria Anthropic uses to choose a model

Anthropic's own guide to choosing the right model weighs three things: capabilities, speed and cost. Capabilities means how well the model handles hard reasoning, long tasks and complex instructions. Speed means how quickly it answers. Cost means what you pay per token.

A token is a piece of text the model processes. Anthropic's pricing page puts one token at roughly 4 characters or 0.75 words in English. Prices are quoted in US dollars per million tokens, with separate rates for input (what you send) and output (what the model writes back).

The three criteria pull against each other. The most capable model is the slowest and the most expensive. The fastest and cheapest model handles simpler work very well but has less headroom for difficult judgement. Every choice is a trade, and the right trade depends on the step.

All current Claude models share a common base. According to Anthropic's models overview, every current model accepts text and images, writes text, works in several languages and can use tools. So the choice is rarely about whether a model can do something at all. It is about how well, how fast and at what price.

Effort settings are often a better lever than switching models

The effort parameter is a setting that trades intelligence against speed and cost within a single model. Anthropic states plainly that "tuning effort is often a better lever than switching models." Before you move a task to a bigger model, try raising effort on the one you have. Before you move down, try lowering it.

Each model starts at a different default. The models overview lists these defaults:

  • Claude Fable 5.1 defaults to high effort.
  • Claude Opus 5.5 defaults to medium effort.
  • Claude Sonnet 5 defaults to high effort.
  • Claude Haiku 4.5 does not support the effort setting.

The Opus 5.5 default matters most. Because it starts at medium, you have room to raise it for a hard case without changing model. Anthropic also notes that Opus 5.5 is no longer available with thinking switched off, and that its adaptive thinking is always on. Adaptive thinking means the model decides for itself how much to reason before it answers.

Haiku 4.5 works differently. It uses manual extended thinking, which you switch on in the request, rather than adaptive thinking. If a Haiku step needs more care, turning on extended thinking is your first lever.

The four current Claude models side by side: price, context and effort

The current lineup has four models open to all customers. The table below compares them on the figures Anthropic publishes on its pricing page and models overview. Context window means how much text the model can read in one request.

ModelInput / output per million tokensContext windowMax outputDefault effort
Claude Fable 5.1$10 / $501M tokens128K tokensHigh
Claude Opus 5.5$4 / $201M tokens128K tokensMedium
Claude Sonnet 5$2 / $101M tokens128K tokensHigh
Claude Haiku 4.5$1 / $5200K tokens64K tokensNot supported

Two cost levers apply across the range. The Batch API processes large volumes of requests asynchronously at a 50% discount on both input and output. A prompt cache hit, where the model reuses text it has already processed, costs 10% of the standard input price on most models.

Haiku 4.5 is the current Haiku and is listed as the fastest model. Anthropic says newer Sonnet and Haiku versions will follow in the coming weeks, with no date given. The advice in this guide is by tier, so it holds when the version numbers change.

Two ways to start: efficiency-first with Haiku 4.5 or capability-first with Opus 5.5

Anthropic describes two sensible starting points, and both end in the same place: the cheapest model that does the job well. The difference is which direction you travel.

Efficiency-first

Start on the fast, low-cost model and move up only when you measure a gap. Anthropic's guide says starting with a model like Claude Haiku 4.5 "can be the optimal approach" for this. It suits high-volume, well-defined work such as sorting messages or pulling fields from documents. You pay the least from day one and only spend more where the results prove you need to. The guide to Claude Haiku 4.5 for fast, cheap work covers where it holds up.

Capability-first

Start on the most capable practical model, get the workflow right, then move down or lower effort. Anthropic states that "most workloads start with Claude Opus 5.5," and the models overview recommends it when you are unsure. This suits complex or unfamiliar work, where you first need to know the task can be done at all. The post on which steps need Claude Opus 5.5 explains the move down in more detail.

Pick the direction by how well you already understand the task. Well-defined and repetitive work points to efficiency-first. New or judgement-heavy work points to capability-first.

Both routes end on the cheapest model that does the job well, starting from opposite ends. Efficiency first / Capability first. Start on: Haiku 4.5 / Opus 5.5; Direction of travel: Move up only where you measure a gap / Move down once the results are

Combining models: a cheaper model does the bulk work and a frontier model decides the hard cases

Most real workflows are cheapest when two models share the work. Anthropic calls these multi-model strategies: a lower-cost model is paired with a frontier model so that most tokens are billed at the lower rate. A frontier model is one of the most capable models available, here Opus or Fable.

Anthropic names two common patterns:

  • Executor and advisor: a cheaper model does the work and escalates to a frontier model when it meets a case it cannot settle.
  • Orchestrator and workers: a frontier model plans the job and hands the individual pieces to lower-cost models.

The price gap makes the first pattern attractive. Haiku 4.5 input costs $1 per million tokens against $4 for Opus 5.5, so every routine case the smaller model settles costs a quarter of the Opus rate on input. Only the escalated cases pay the higher price.

At Svennis we usually put the triage step on the smallest model and route only the cases it marks as uncertain to a larger one. We then review a sample of the routed items by hand before we change any model, because the escalation rule matters as much as the model choice.

Decision table: which Claude model to start on for six small-business jobs

The table below turns the guidance into jobs a small business actually runs. The starting model follows Anthropic's matrix: Haiku for real-time, high-volume and sub-agent work, Sonnet for everyday analysis and content, Opus for complex agentic and enterprise work, and Fable for hours-long sessions and deep research.

JobStart onSign to move up or down
Sorting the support inbox into queuesHaikuUp to Sonnet if too many messages land in the wrong queue
Extracting fields from invoicesHaiku, via the Batch APIUp to Sonnet if unusual layouts are misread
Drafting a client proposalSonnetUp to Opus if drafts miss commercial nuance; down to Haiku never for client-facing work
Monthly management report from CRM dataSonnet, via the Batch APIUp to Opus if the analysis misreads trends
An agent working through a backlogOpusDown to Sonnet or lower effort once the steps are stable
A long research taskOpus at higher effortUp to Fable if Opus at higher effort still falls short

Two rows use the Batch API because neither job needs an instant answer. A monthly report or a pile of invoices can wait for asynchronous processing, and the 50% discount applies. The post on Claude Sonnet 5 as the everyday business model covers the drafting and reporting rows in more depth.

Worked example: a support inbox that sorts, drafts and escalates with three models

A support inbox shows model selection at work because every message passes through three different jobs. Take a company that handles customer requests in Zoho Desk and keeps customer records in Zoho CRM.

  1. Sort every message with Haiku 4.5. The model reads each incoming email, assigns a category and picks the right queue. This is high-volume, well-defined work, and Haiku is the fastest and cheapest model. Its 200K token context window is ample for a single email with its history.
  2. Draft routine replies with Sonnet 5. For standard requests, Sonnet writes a reply for an agent to check. Drafting is everyday content work, which Anthropic places with Sonnet.
  3. Escalate hard cases to Opus 5.5. A disputed refund or a complaint that threatens a contract goes to Opus, with the customer's record attached. These cases are few, so the higher price applies to a small share of the traffic.

The split keeps the bill in proportion to the difficulty of the work. Most tokens flow through the cheapest model, and the expensive model sees only the messages that justify it.

Routing accuracy decides whether the whole design works. A message in the wrong queue waits longer and costs more to fix. The guide to AI email triage and first-time-right routing covers how to measure it.

Where Claude Fable 5.1 fits at the top end

Claude Fable 5.1 is Anthropic's most capable model open to all customers, and most businesses will need it rarely. Anthropic recommends it for demanding reasoning and long-horizon agentic work, meaning agent sessions that run for hours. It also recommends Fable "when your evals on Claude Opus 5.5 at higher effort still fall short."

That second condition is the practical test. Anthropic's newsroom says Opus 5.5 performs at the level of Fable 5.1 on most work. So the sensible order is to try Opus 5.5 first, raise its effort, and move to Fable only when measured results still fall short. At $10 input and $50 output per million tokens, Fable costs two and a half times the Opus 5.5 rate.

A related model, Claude Mythos 5.1, offers the same capabilities as Fable 5.1 but is available only to participants in Anthropic's Project Glasswing. For a normal business account, Fable is the top of the range. The post on Claude Fable 5.1 and its limits covers where it helps and where it does not.

At max effort Opus 5.5 scored 1846 Elo on GDPval, above Fable 5.1 at 1735: Opus 5.5 at max effort 1846, Fable 5.1 1735, Opus 5 1708 (Elo)
Source: anthropic.com

Choosing the Claude model in the Claude apps and on the API

How you choose a model depends on how you use Claude. In the Claude apps on a paid plan, you pick the model for each conversation. On the API, the model is a model id you set in each call your system makes.

The app choice is a habit, not a configuration. Encourage staff to use a smaller model for quick questions and a larger one for a proposal or a tricky analysis. Anthropic has raised five-hour usage limits on the Pro, Max, Team and seat-based Enterprise plans, but the principle of matching model to task still stretches those limits further.

The API choice is part of the system design. Anthropic's docs give ids such as claude-opus-5-5, claude-fable-5-1 and claude-haiku-4-5. The Haiku alias resolves to a pinned snapshot, claude-haiku-4-5-20251001, so its behaviour does not shift under you.

Each cloud platform uses its own ids. Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS each list separate model ids in the models overview. If your system runs through one of them, take the id from that platform's column. The Models API can also report each model's capabilities and token limits programmatically.

Data location for UK and EU companies using Claude models

On Anthropic's own API, stored data sits in the United States. The data residency documentation separates two settings. Workspace geo controls where data is stored at rest, and "us" is currently the only option. Inference geo controls where the model runs each request.

Anthropic offers no EU or UK option for either setting. The default inference geo is "global", meaning inference may run in any available geography. On the newer models you can restrict inference to "us", which costs 1.1 times the standard rate. Haiku 4.5 does not accept the inference geo parameter at all, and a request that includes it returns a 400 error.

If you need processing in a particular region, look at the cloud platforms. Claude models are available through Amazon Bedrock, Google Cloud and Microsoft Foundry, and Bedrock and Google Cloud set their own regional pricing. Check the specific model and region on that provider's own pages before you design around it, because availability differs by model.

Two further points help a European compliance review. Anthropic says Opus 5.5 is available with zero data retention and includes watermarking measures to comply with the EU AI Act. Model choice and data location are therefore linked: the region you need may narrow the models you can use.

Next steps: test twenty real examples on two Claude models before you commit

Anthropic calls a use-case-specific evaluation set "the most important step" in choosing a model. Its Opus 5.5 announcement adds that benchmark margins have become a less reliable guide to real-world differences. Your own examples beat any published score.

A simple test takes a few steps:

  1. Pick one workflow step, such as sorting inbox messages or drafting proposals.
  2. Collect twenty real examples, including a few awkward ones.
  3. Run all twenty on two neighbouring models, for example Haiku and Sonnet.
  4. Compare the answers side by side and mark each as usable or not.
  5. Compare the bill for each run in the Claude Console.

Choose the cheaper model if it gets the answers right. If it falls short, try raising effort or turning on extended thinking before you switch. Repeat the test when a new version arrives, since the tiers stay stable even as the numbers change.

If you want to see how these choices fit into systems you already run, the guide to putting Claude to work inside your existing tools is the natural next read.

Sources

  1. 1. Choosing the right model - Claude Platform Docs
  2. 2. Models overview - Claude Platform Docs
  3. 3. Models overview - Claude Platform Docs (docs.anthropic.com)
  4. 4. Pricing - Claude Platform Docs
  5. 5. Claude Haiku 4.5 - Claude Platform Docs
  6. 6. Introducing Claude Opus 5.5 - Anthropic
  7. 7. Newsroom - Anthropic
  8. 8. Data residency - Claude Platform Docs

Related articles