Svennis AI
10 min read

Moving from Claude Opus 5 to Opus 5.5 in live business systems

Opus 5.5 costs less per token than Opus 5, but four API changes can break a live integration. This guide covers the prices, the fixes and a checklist to run before you switch.

Abstract cover of two parallel paths converging into a single lighter, faster line after a small step

Moving from Claude Opus 5 to Opus 5.5: change one ID, then fix four settings

Moving from Claude Opus 5 to Opus 5.5 takes one line of code and a round of careful testing. You change the model ID from claude-opus-5 to claude-opus-5-5. Then you remove the settings that Opus 5.5 rejects. The reward is a lower bill: Opus 5.5 costs $4 and $20 per million input and output tokens, against $5 and $25 for Opus 5.

A model migration is the switch of a live system's API calls from one model to another, together with the tests that prove prompts, routing and integrations still behave. On Amazon Bedrock the IDs are anthropic.claude-opus-5 and anthropic.claude-opus-5-5. The other platforms use the same IDs as the Claude API.

Anthropic calls Opus 5.5 "the current Opus model". Its models overview says that if you are unsure which model to use, you should start with Opus 5.5 for most workloads. Both models have a 1M-token context window and a maximum output of 128K tokens, so long documents and long answers fit on both.

This guide covers the price difference, where Anthropic's 40% saving comes from, the four breaking changes, the quieter behaviour changes, the timeline, a worked cost example and a checklist. It is written for a business that already runs something on Opus 5 and wants to switch without surprises.

Opus 5.5 prices are 20% lower per token, and cache reads are 60% lower

Every Opus 5.5 token price is lower than the matching Opus 5 price. Input and output tokens cost 20% less, and cache reads cost 60% less. All prices on Anthropic's Claude pricing page are in US dollars.

Price per million tokensOpus 5Opus 5.5Change
Input$5$420% less
Output$25$2020% less
Cache read$0.50$0.2060% less
5-minute cache write$6.25$520% less
1-hour cache write$10$820% less
Batch input / output$2.50 / $12.50$2 / $1020% less

Prompt caching is a feature that reuses parts of a prompt already processed in earlier calls, so the model reads them from cache at a fraction of the input price. On Opus 5.5 a cache read costs 0.05 times the base input price. Anthropic says cache reads "make up the majority of agentic and coding work costs". That is why the 60% cut matters more than the headline 20% for systems that resend the same long instructions on every call.

The Batch API is Anthropic's route for processing large volumes of requests without waiting for each answer. It halves both input and output prices on both models. The minimum cacheable prompt is 512 tokens on both, so your caching setup carries over unchanged.

Opus 5.5 tokens cost 20% less than Opus 5, and cache reads cost 60% less: Opus 5.5 input and output tokens 20% cheaper than Opus 5, Opus 5.5 cache reads 60% cheaper than Opus 5, Notice before a public model is retired 60 days at least
Source: anthropic.com, platform.claude.com

Anthropic's 40% saving depends partly on the lower default effort

Anthropic reports a larger saving than the per-token cut. Its announcement says: "at default settings it will cost 40% less than Opus 5 on typical workloads." The per-token price explains 20% of that. The rest comes from how many tokens the model uses to finish a task.

One default has changed, and it affects token use. Effort is the control for thinking depth, latency and cost. On Opus 5 a request that did not set effort ran at high. On Opus 5.5 the same request runs at medium. Less thinking usually means fewer output tokens per task, and output tokens are the expensive ones.

This has two consequences for a live system.

  • If your code never sets effort, you get the lower default and most of the saving, but the depth of reasoning also changes.
  • If you set effort to high to keep Opus 5 behaviour, expect a saving closer to the 20% per-token cut than to 40%.

Neither choice is right for every step. A classification step may be fine at medium. A step that drafts a contract clause may need more. Test each step at medium first and raise effort only where the results fall short.

Some steps may not need Opus at all. If a step is simple and high-volume, compare it against the smaller models before you pay Opus prices for it. Our guide on matching Opus, Sonnet or Haiku to each business task covers that decision.

Opus 5.5 quality and speed against Opus 5, by Anthropic's own measurements

Anthropic says Opus 5.5 "performs at the level of Claude Fable 5.1 on most work". That is notable on price alone. Claude Fable 5.1 costs $10 and $50 per million input and output tokens, two and a half times the Opus 5.5 rate.

Speed improves too. Anthropic states that Opus 5.5 "generates output more than 30% faster than Opus 5". For a service desk or a chat assistant, faster output means a shorter wait before the user sees an answer.

Anthropic's own benchmark table compares only Claude models. On GDPval-AA v2.1, with Opus 5.5 at max effort, Anthropic reports these results:

  • Opus 5.5: 1846 Elo
  • Claude Fable 5.1: 1735 Elo
  • Opus 5: 1708 Elo

Treat these as Anthropic's measurements at max effort, not at the medium default your system will use. Anthropic itself cautions that "benchmark margins have become a less reliable guide to real-world differences". Your own cases are the only benchmark that decides the switch.

The reliable knowledge cutoff moves from May 2026 on Opus 5 to June 2026 on Opus 5.5. For a wider view of where Opus 5.5 earns its place in a business process, see our guide to which business steps need Claude Opus 5.5. This post stays with the migration.

Four breaking changes that can stop an Opus 5 integration

Anthropic lists four breaking changes for code already running on Opus 5. Each one returns an error or changes a request, so an unchanged integration can fail on the first call. The details are on Anthropic's What's new in Claude Opus 5.5 page.

1. Thinking can no longer be turned off

Adaptive thinking is always on in Opus 5.5. A request with thinking: {"type": "disabled"} returns a 400 invalid_request_error. So does a manual budget set with budget_tokens. Remove both and control depth with the effort parameter instead.

2. Forced tool use is not supported

Setting tool_choice to {"type": "any"} or to a named tool returns a 400 error. Use {"type": "auto"}. Routing systems often force a tool call to classify a ticket, so check this first. If you route requests from chat, as in our guide to Claude in Slack for internal support, make the prompt ask for the tool and have your code check that a tool call came back.

3. Thinking blocks are tied to the model and the conversation

Opus 5.5 reads thinking blocks from Opus 5 and earlier Opus, Sonnet and Haiku models. It does not read those from Claude Fable or Claude Mythos models. The API drops unreadable blocks, the request succeeds and dropped blocks are not billed. For accounts created on or after 31 August 2026, the API rejects replayed thinking blocks with a 400 error if earlier context was edited.

4. The earlier computer-use tool version is rejected

On the Claude API and Google Cloud, Opus 5.5 rejects the earlier computer-use tool version that Opus 5 accepted. On Amazon Bedrock it still works, so no change is needed there.

Disabled thinking and forced tool use return errors on Opus 5.5, while default effort drops to medium. Opus 5.5 behaviour / What to change. Thinking set to disabled: Request rejected with an error / Remove it and set effort instead; Manual budget_tok

Default effort, progress updates and refusals change behaviour without an error

Three Opus 5.5 changes produce no error at all. They are harder to catch than the breaking changes, because the system keeps running and simply behaves differently.

The default effort is medium. A request that omits effort runs at medium on Opus 5.5, where it ran at high on Opus 5. Answers may be shorter or less thorough on hard cases. Compare outputs on difficult examples, not just easy ones.

Progress updates go quiet. On Opus 5.5, the text the model writes between tool calls arrives as thinking blocks rather than text blocks. At the default display setting of "omitted", the text in those blocks is empty. An application that streams progress messages to its users goes quiet between tool calls, with no error.

Anthropic's fix for quiet progress updates is to set a display value that returns the text. If your assistant shows "checking your order" style messages, test this before you switch.

Refusals return HTTP 200. When Opus 5.5 declines a request, the response has stop_reason: "refusal" and a stop_details object naming the policy area. The status code is still 200. If your integration only checks the status code, it will pass a refusal downstream as if it were an answer. Add a check on stop_reason so a refusal goes to a person instead.

None of these three shows up in an error log. Only a side-by-side comparison of real outputs will reveal them.

Opus 5 remains available until at least 24 July 2027, so you can test first

You do not have to switch this week. Anthropic released Opus 5 on 24 July 2026 and lists its status as "Active (legacy)", with retirement not sooner than 24 July 2027. Legacy means the model will no longer receive updates and may be deprecated in the future.

Anthropic's deprecation policy gives some protection. For publicly released models, Anthropic provides at least 60 days' notice before retirement to customers with active deployments. After retirement, requests to the model fail. A system still pointing at claude-opus-5 on that day stops working.

The dates depend on where you call the model. Anthropic's dates apply to its own platforms: the Claude API, Claude Platform on AWS and Microsoft Foundry. Amazon Bedrock and Google Cloud set their own retirement schedules, so the status and dates there can differ. If you use one of those, check its own lifecycle page.

Opus 5.5 was released on 22 September 2026, with retirement not sooner than 22 September 2027. That gives a system moved now about a year on the newer model before its own retirement date could arrive. A calm plan is to test in the next few weeks, switch one workload at a time, and keep Opus 5 as a fallback until the numbers settle.

A worked example: one monthly ticket-routing workload priced on both models

A worked example shows what the price change means for a real bill. The workload below is illustrative, not measured at any client. It is a ticket-routing assistant that reads incoming requests and files them in Zoho Desk. Its long system prompt, with the routing rules and team list, is cached.

Assume one month of this workload uses 20 million input tokens, 4 million output tokens and 50 million cache-read tokens. Here is the cost at Anthropic's published prices, with the same token counts on both models.

LineTokens per monthOpus 5Opus 5.5
Input20 million$100$80
Output4 million$100$80
Cache reads50 million$25$10
Total$225$170

With identical token counts, the monthly bill falls from $225 to $170, about 24% lower. The cache line falls furthest, which is why cached prompts gain most from the switch.

The real saving can be larger. At the medium default effort, Opus 5.5 may use fewer output tokens per ticket. If Anthropic's 40% figure held for this workload, the bill would be about $135. Only your own usage data will show which number applies. Run the same month of tickets on both models and read the usage export before you believe either figure.

What the switch means for a UK company: dollar prices, watermarks and platforms

For a UK company, three practical points follow from Anthropic's pages. None of them blocks the switch, but each belongs in your notes.

Prices are in US dollars. Anthropic bills in USD, so budget in dollars and let your finance team convert. On Claude Platform on AWS and Microsoft Foundry, usage is converted to Claude Consumption Units at $0.01 each. US-only inference carries a 1.1 times price multiplier on current models. A UK business has no reason to request it by default.

Output carries a watermark. Since 2 August 2026, the EU requires AI providers serving its market to mark AI-generated content. Anthropic says Opus 5.5 "comes with our watermarking measures to comply with the EU AI Act". Anthropic applies watermarking globally because it cannot yet scope it by region, so UK output carries it too.

The watermark adds no text, no hidden characters and no extra tokens, so it adds nothing to your bill. It also carries nothing that identifies a person, an organisation or a chat. For the rules that apply to your own use of AI, see our overview of AI law in the UK and what applies to your business.

Platform choice changes the dates. Opus 5.5 is available on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. If you run through Bedrock or Google Cloud, their retirement schedule for Opus 5 may differ from Anthropic's. Fast mode, a lower-latency option in research preview at $8 and $40 per million tokens, runs on the Claude API only.

Migration checklist: what to change and what to test before switching

The checklist below covers every change in this guide, in the order to do them. Work through it for each system that calls Opus 5.

StepWhat to doHow you know it worked
Find every callerExport usage from the Usage page in Claude ConsoleThe CSV lists each API key calling Opus 5
Change the model IDclaude-opus-5-5, or anthropic.claude-opus-5-5 on BedrockRequests reach the new model
Remove disabled thinkingDelete "disabled" and any budget_tokensNo 400 errors
Replace forced tool useSet tool_choice to auto, check for the tool callEvery case still returns a tool call
Check computer useUpdate the tool version on the Claude API and Google CloudNo rejected tool errors
Decide effort per stepStart at medium, raise only where results fall shortQuality matches Opus 5 on hard cases
Fix progress displaySet a display value that returns textUsers see updates between tool calls
Handle refusalsCheck stop_reason for "refusal"Refusals reach a person
Rerun twenty real casesSame inputs on both modelsRouting and answers match or improve
Compare the billRead a week of usage on bothCost per case is lower

The comparison step is where most problems appear. At Svennis we run the old and new model side by side on the same set of real tickets and compare the routing decisions line by line before anything changes in production. A difference in one decision matters more than an improvement in the average.

Next steps: audit your usage, test on real cases, then switch one workload

The first step is to find out what you actually run on Opus 5. Export the usage CSV from Claude Console and list each API key and the system behind it. Most businesses find one or two workloads that account for most of the spend.

Then take these steps for the largest workload first.

  1. Pick twenty real cases from last month, including the awkward ones.
  2. Make the code changes from the checklist in a test copy of the integration.
  3. Run the twenty cases on both models at the default effort.
  4. Compare the outputs case by case, then compare the token counts.
  5. Switch that workload, keep Opus 5 as a fallback for a week, then move to the next.

Because Opus 5 stays available until at least 24 July 2027, you have time to do this properly. There is no reason to switch everything on the same day.

If your Opus 5 system is a service desk or a routing assistant, the cases worth testing look much like the ones in our guide to a Claude service desk in Teams that routes into Zoho Desk. Use it to decide which cases belong in your test set.

Sources

  1. 1. Anthropic: Introducing Claude Opus 5.5
  2. 2. Anthropic Newsroom
  3. 3. Claude Platform Docs: What's new in Claude Opus 5.5
  4. 4. Claude Platform Docs: Claude Opus 5.5
  5. 5. Claude Platform Docs: Claude Opus 5
  6. 6. Claude Platform Docs: Pricing
  7. 7. Claude Platform Docs: Model deprecations
  8. 8. Claude Platform Docs: Models overview
  9. 9. Anthropic: How Claude's text watermarking works

Related articles