Svennis AI
10 min read

Claude Sonnet 5 as the everyday business model and when to choose something bigger

Claude Sonnet 5 suits most routine business work at $2 per million input tokens. This guide covers when to use it, what changes if you migrate, and where your data goes.

Abstract layered shapes of different sizes, with a mid-sized form carrying most of the flowing movement

Claude Sonnet 5 is the everyday business model for routine work

Claude Sonnet 5 as the everyday business model makes sense for most routine work: ticket routing, CRM updates and first drafts. It costs half as much per token as Opus 5.5 and a fifth as much as Fable 5.1. Anthropic says its performance is close to that of a recent Opus model. Keep a larger model for the few steps that fail on Sonnet 5.

Claude Sonnet 5 is Anthropic's mid-tier model, released on 30 June 2026. It is the default model on the Free and Pro plans, and Max, Team and Enterprise users can choose it too. Developers call it through the Claude API as claude-sonnet-5.

This post covers Sonnet 5, the everyday model, in a set of three on the current Claude models. Claude Opus 5.5 sits one step up, and our guide to when Opus 5.5 is worth using covers which steps need it. Claude Fable 5.1 sits at the top, for demanding reasoning and long-horizon agentic work, and has its own post.

The rest of this guide gives you the facts, a decision table, a worked example, the migration changes and the data residency options for UK and EU businesses.

Sonnet 5 facts: prices, context window and limits against the other models

Sonnet 5 costs $2 per million input tokens and $10 per million output tokens. A token is a piece of text the model processes, roughly four characters or three quarters of a word in English. Anthropic first announced these as introductory prices, then made them permanent. The $3 input and $15 output pricing once due on 1 September no longer applies.

The table compares the current models on Anthropic's own pricing and models pages. All prices are in US dollars.

ModelInput / output per million tokensContext windowMax outputDefault effortKnowledge cutoff
Claude Sonnet 5$2 / $101M tokens128K tokenshighJanuary 2026
Claude Opus 5.5$4 / $201M tokens128K tokensmediumJune 2026
Claude Fable 5.1$10 / $501M tokens128K tokenshighJune 2026
Claude Haiku 4.5$1 / $5200K tokens64K tokensnot supportedFebruary 2025

Sonnet 5 takes text and images as input and produces text. It runs on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic's pricing page describes it as "the best combination of speed and intelligence". Its retirement date is not sooner than 30 June 2027.

Sonnet 5 is also the first Sonnet-tier model with real-time cybersecurity safeguards. It launched with these cyber safeguards on by default. They detect and block dangerous cyber usage as it happens.

Why ticket routing, CRM updates and drafting fit Sonnet 5

Routine business tasks fit Sonnet 5 because they are short, repetitive and easy to check. Routing a ticket means reading a few paragraphs and picking a queue. Updating a Zoho CRM record means pulling a name, a date or an order number out of an email. Drafting a reply means following a template your team already uses.

None of these steps needs the deepest reasoning available. They need accuracy on a narrow job, low cost per request and a quick answer. Sonnet 5 meets all three at $2 per million input tokens.

Anthropic's safety assessments also matter for this kind of work. Sonnet 5 shows lower rates of hallucination and sycophancy than the previous Sonnet. Hallucination is when a model states something false as fact. Sycophancy is when it agrees with the user rather than the evidence. Both cause silent errors in CRM fields and routed tickets.

Anthropic's own models page takes a different starting point. It says that if you are unsure, you should start with Opus 5.5 for most workloads. That advice suits a team with no test data. For a routine task you can measure, starting on Sonnet 5 and moving up only where tests fail costs less. Our post on Claude in Slack routing internal support tickets shows the same pattern on a different front end.

Signs a task needs Opus 5.5 or Fable 5.1 instead of Sonnet 5

A task needs a larger model when Sonnet 5 keeps failing the same test after you have fixed the prompt. Model choice should follow evidence from your own examples, not the length or importance of the task. At Svennis we build a small set of real, anonymised tickets or emails before choosing a model, and we move a step up to Opus 5.5 only when it keeps failing that set at high effort.

These signs point towards Opus 5.5:

  • The step runs for a long time across many tools. Anthropic describes Opus 5.5 as built "for long-running agentic coding and knowledge work".
  • The answer depends on the model's own knowledge of recent events. Sonnet 5's reliable knowledge cutoff is January 2026, against June 2026 for Opus 5.5.
  • Errors repeat in the same place, such as a judgement call your rules cannot capture.

Opus 5.5 costs twice as much per token as Sonnet 5, at $4 input and $20 output. Anthropic says Opus 5.5 performs at the level of Fable 5.1 on most work. It also says Opus 5.5 costs 40% less to run than Opus 5.

Fable 5.1 is the next step after that. Anthropic recommends it for demanding reasoning and long-horizon agentic work, or when your evals on Opus 5.5 at higher effort still fall short. At $10 input and $50 output, it belongs in few small-business workflows.

Model decision table: which Claude model to start each business task on

A model decision table gives each task a starting model and a clear trigger for moving up. The table below reflects the capabilities and prices on Anthropic's pages. The triggers are the tests you run on your own data.

TaskStart onMove up when
Routing tickets to the right queueSonnet 5The same category errors persist after prompt fixes
Updating CRM fields from emailsSonnet 5Fields need judgement your rules cannot describe
Drafting replies and quotes from templatesSonnet 5Staff rewrite most drafts rather than edit them
Overnight summaries of many recordsSonnet 5 on the Batch APISummaries miss links between records
Very high volumes of simple, short classificationHaiku 4.5, if tests passInputs exceed 200K tokens or need effort control
Multi-step agents working across several systemsOpus 5.5Evals at higher effort still fall short: try Fable 5.1

Haiku 4.5 costs $1 input and $5 output, half the price of Sonnet 5. It has a 200K-token context window, 64K tokens of output and no effort setting. Its retirement date is not sooner than 15 October 2026, so plan any Haiku 4.5 workflow with that date in mind.

The Batch API processes large volumes of requests asynchronously at a 50% discount on input and output tokens. On Sonnet 5 that means $1 input and $5 output, the same as Haiku 4.5.

Worked example: routing Zoho Desk tickets on Sonnet 5

A service desk that routes Zoho Desk tickets is a typical Sonnet 5 job. Staff raise a request, the model reads it, picks a department and writes the ticket. Our post on a Teams service desk in front of Zoho Desk covers the full build. These are the model settings for the routing step:

  1. Model. Call claude-sonnet-5 on the Claude API, or the Sonnet 5 model ID on your cloud platform.
  2. Fixed instructions first. Put the department list and routing rules at the start of every prompt and cache them. A cache read costs $0.20 per million tokens, against $2 for normal input.
  3. Effort. Effort controls how deeply the model thinks, and it defaults to high. Test a lower setting against your ticket set before changing it.
  4. No sampling parameters. Leave out temperature, top_p and top_k. Non-default values return a 400 error.
  5. Output limit. Set max_tokens high enough for thinking plus the answer, because it caps both.
  6. Refusals. Check stop_reason. A refusal comes back as a successful HTTP 200 response with stop_reason: "refusal", not as an error.
  7. Write-back and log. Create the ticket in the chosen department and log the model's choice for review.

The refusal check is the step teams most often miss. Code that only checks the HTTP status will treat a refusal as a routed ticket.

What Sonnet 5 costs in practice, and why token counts rise

Sonnet 5 uses a new tokenizer, so the same text produces more tokens than before. A tokenizer is the part of the system that splits text into tokens. Anthropic's migration guide says the same input produces approximately 30% more tokens than on the previous Sonnet. Its announcement puts the range at roughly 1.0 to 1.35 times, depending on the content type.

The per-token price is lower, at $2 and $10 against the previous Sonnet's $3 and $15. Anthropic states plainly that the cost of an equivalent request does not drop in direct proportion. Measure your own requests before you forecast savings.

Three more factors shape your bill:

  • Thinking tokens. They are billed as output tokens even when the thinking text is not returned. On Sonnet 5, thinking.display defaults to omitted.
  • Caching and batching. A cache hit costs 10% of the input price. The Batch API halves input and output prices.
  • US-only inference. Pinning inference to the US costs 1.1 times the standard rate across all token categories.

The Usage page in Claude Console exports a CSV broken down by API key and model. Run a week of real traffic on Sonnet 5 and read that file before you commit to a monthly budget.

Sonnet 5 input costs $2 against $3 before, yet the same text makes about 30% more tokens: Sonnet 5 input price 2 USD per million tokens, Previous Sonnet input price 3 USD per million tokens, Extra tokens for the same text on Sonnet 5 30% more tokens
Source: platform.claude.com

Migrating from the previous Sonnet: three behaviour changes to fix

Sonnet 5 is a drop-in upgrade for the previous Sonnet, but three behaviours change. If you have an integration built, check each one before switching the model ID.

  1. Adaptive thinking is on by default. Adaptive thinking lets the model decide how much to think, guided by the effort setting. Expect more output tokens at the default high effort.
  2. Manual extended thinking is rejected. A request with thinking: {type: "enabled", budget_tokens: N} returns a 400 error. Remove it and use the effort parameter instead.
  3. Non-default sampling parameters are rejected. Setting temperature, top_p or top_k to a non-default value returns a 400 error. In the Python SDK from v1.0, passing them raises a TypeError.

Some other details carry over or change quietly. Prefilling the assistant message still returns a 400 error, as it did on the previous Sonnet. Priority Tier is not available on Sonnet 5. Thinking text is omitted by default, so any code that reads a thinking summary will find an empty field.

Sonnet 5 also adds tools. On the Claude API and Google Cloud it supports computer use, as the stable computer_toolset_20260801 toolset, and the browser use tool. The previous Sonnet supports neither.

Three settings change on Sonnet 5, and two older ones return an error until you remove them. What happens on Sonnet 5 / What to do before switching. 1. Adaptive thinking: On by default and guided by the effort setting / Budget for more output tokens

Still on an older Sonnet release? Plan the move before 29 September 2026

An older Sonnet release, the one before the previous Sonnet, has a retirement date of not sooner than 29 September 2026 on Anthropic's deprecations page. Once a model is retired, requests to it fail. Anthropic gives at least 60 days' notice to customers with active deployments, but you should not wait for the email.

Take these steps now if any of your systems still call that release:

  1. Export the usage CSV from Claude Console and find every API key that calls the older model.
  2. Move each workflow to claude-sonnet-5, applying the three behaviour changes above.
  3. Rerun your test set, because the new tokenizer and adaptive thinking change both cost and output.
  4. Check your cloud platform's own schedule if you run on Amazon Bedrock or Google Cloud.

Anthropic's dates apply to the Claude API, Claude Platform on AWS and Microsoft Foundry. Bedrock and Google Cloud set their own retirement schedules, so dates there can differ.

Moving also gives you control over where inference runs. The older release returns a 400 error if a request includes the inference_geo parameter. The previous Sonnet stays active, with retirement not sooner than 17 February 2027, but Sonnet 5 is the better target.

Where Sonnet 5 data goes: UK and EU residency options

Anthropic's own API stores your data in the US and offers no EU or UK residency itself. Workspace geo, the setting for where data is stored at rest, currently has only one value: "us". It is set when you create a workspace and cannot be changed afterwards.

Inference geo, the setting for where the model runs, defaults to "global". That means inference may run in any available geography. You can pin it to "us" at 1.1 times the standard rate, but there is no EU or UK option.

EU residency for Sonnet 5 comes only through Amazon Bedrock, Google Cloud or Microsoft Foundry. The Amazon Bedrock regional availability page sets out three routing options:

  • In-Region: requests never leave the AWS Region you specify.
  • Geographic: requests stay within a defined geography, such as the EU.
  • Global: requests go to a supported commercial Region anywhere in the world.

On Bedrock, Sonnet 5 is available in-Region in a single EU Region only: Ireland. Elsewhere in the EU you reach it through the EU cross-Region profile or the global endpoint. No Claude model runs in-Region in AWS London. Cross-Region inference is priced at source Region rates with no routing surcharge.

Google Cloud and Microsoft Foundry set their own regional options and pricing, so check them in your console before you commit. Our GDPR checklist for Claude covers the contract side.

What Sonnet 5 means for a UK or EU business

For most UK and EU businesses, Sonnet 5 matches the AI work they actually do: handling text. The UK government's AI adoption research surveyed 3,500 businesses between February and May 2025. Around 1 in 6 (16%) used at least one AI technology. Among those users, 85% used natural language processing and text generation.

The same research found that 67% of AI users report significant human input or checking of AI outputs. Keep that habit with Sonnet 5. Lower hallucination rates reduce errors; they do not remove them. The international scientific report on AI safety states that current techniques for explaining why a model produces a given output are severely limited.

Data location is the second point for UK firms. Because no Claude model runs in-Region in AWS London, your prompts leave the UK on every route. The closest single-Region option for Sonnet 5 on Bedrock is Ireland. Decide whether that fits your obligations before you connect customer records. Our overview of AI law in the UK for businesses sets out the rules that apply.

Micro businesses use AI least, at 14%, against 36% of large firms. The most cited reasons are a lack of identified need and limited skills. A single routed queue on Sonnet 5 answers both: one clear need, one measurable result.

Next steps: test one routine task on Sonnet 5

The practical next step is to pick one routine task and prove it on Sonnet 5 before choosing anything bigger. Ticket routing, CRM field updates and template drafts are the usual candidates.

  1. Collect a set of real, anonymised examples with the correct answer for each.
  2. Decide where the data may go: Anthropic's API in the US, or Bedrock in Ireland or the EU profile.
  3. Run the set on claude-sonnet-5 with cached instructions and no sampling parameters.
  4. Check stop_reason on every response and log each result.
  5. Export the usage CSV after a week and compare cost with the old process.
  6. Move only the failing steps to Opus 5.5, and retest.

If you still call an older Sonnet release, fix that first, before 29 September 2026. Then read our guide to putting Claude inside the tools you already use to plan where the tested task should live.

Sources

  1. 1. Amazon Bedrock: Regional availability by models
  2. 2. Claude Platform Docs: Data residency
  3. 3. Claude Platform Docs: Pricing
  4. 4. Claude Platform Docs: Model deprecations
  5. 5. Claude Platform Docs: Migrating to Claude Sonnet 5
  6. 6. Claude Platform Docs: Claude Sonnet 5 overview
  7. 7. Anthropic: Introducing Claude Sonnet 5
  8. 8. Anthropic: Newsroom
  9. 9. Claude Platform Docs: Models overview
  10. 10. GOV.UK (DSIT): AI Adoption Research
  11. 11. GOV.UK (DSIT): International scientific report on the safety of advanced AI, interim report

Related articles