In the last week, Anthropic released two new Claude models. If you are using Claude Code or the Claude API you now see new models to choose from in the picker and new default settings. Developers also face API changes that can break code written for the previous models so that's something to be aware of.
Claude Opus 5.5 arrived about a week ago and Claude Sonnet 5.5 arrived yesterday, the 28 September. These are the first two models in Anthropic's Claude 5.5 family. Anthropic describes Sonnet 5.5 as the faster, lower-cost model for well-scoped work, and Opus 5.5 as the model for complex work that needs careful judgement.
The sections below cover what each model costs and how the two compare on Anthropic's benchmarks. Later sections deal with Claude Code, API integrations, safeguards and the limitations Anthropic documents for each model. Every figure comes from Anthropic's announcements and platform documentation, checked on 29 September 2026.
Contents
The Claude 5.5 Family So Far
Claude Sonnet 5.5
Claude Opus 5.5
Sonnet 5.5 and Opus 5.5 Side by Side
Using the New Models in Claude Code and the Claude Apps
What Changes for API Integrations
Safeguards and Refusals
Limitations
Summary
Sources
The Claude 5.5 Family So Far
Anthropic's current lineup has four models. Claude Fable 5.1 is for demanding reasoning and long-horizon agentic work, Claude Opus 5.5 is for long-running agentic coding and knowledge work, and Sonnet 5.5 is "the best combination of speed and intelligence". Haiku 4.5 (which I never use) is the fastest model in the lineup.
The Sonnet 5.5 announcement says that Haiku 5.5 will join the family "in the coming weeks". Until then, Haiku 4.5 remains the current lowest level intelligence model. Sonnet 5 moves to Anthropic's list of legacy models, which are still available in your model picker. Expect this to disappear over the next six months, I'd say.
The three newest models, Fable 5.1 and the two 5.5 releases share the same core specifications. Each has a context window of one million tokens, which Anthropic puts at roughly 555,000 words. Each can also return up to 128,000 tokens in a single response, and Anthropic gives June 2026 as the reliable knowledge cutoff for all three. In other words, don't rely on any of these models to be up to speed with current events. Instead, use web fetch tools for up-to-date news and information on any given topic. The July update on Claude Opus 5 covers the model that Opus 5.5 replaces.
Claude Sonnet 5.5
Anthropic positions Sonnet 5.5 as "a faster, lower-cost complement to Claude Opus 5.5". According to its announcement, the model is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides and spreadsheets. Anthropic adds that it has "a sharp eye for design".
For those of you using Claude models through the API, the price per token is the same as for Sonnet 5. Input costs $2 per million tokens and output costs $10, with cache reads at $0.20 per million. Anthropic reports that Sonnet 5.5 needs far fewer tokens for the same work. In its testing, that brought the cost per task down by up to 30% compared with Sonnet 5. According to Anthropic, the model also generates output more than 30% faster than Sonnet 5, which Anthropic says makes it the fastest Sonnet model to date.
If you access these models through a subscription (Pro, Max, Team, Enterprise) then cost per token is not a concern for you. Your cost is tokens per session and tokens per conversation. Limits reset every five hours, and if you turn on usage credits to keep working past them, that extra usage is billed at standard API rates.
Anthropic published the following benchmark results for Sonnet 5.5 against Sonnet 5.
Terminal-Bench 4.0. On this agentic coding test, Sonnet 5.5 scores 70.6% against 10.3% for Sonnet 5.
CursorBench 4.0. The coding score rises to 55.5% from 34.1%.
FrontierCode 1.1, Main. At Max effort, Sonnet 5.5 reaches 46.2% against 42.4% for its predecessor.
GDPval-AA v2.1. On this test of real-world work across occupations, the score rises to 1844 from 1449.
OSWorld 2.1, partial. The computer use result moves to 80.1% from 57.0%.
Humanity's Last Exam, with tools. Sonnet 5.5 scores 64.5% against 54.9% for Sonnet 5.
Anthropic also reports that Sonnet 5.5 is strong on long-horizon work and image understanding. The company states that it writes more clearly than the previous generation of models, and like Opus 5.5 and Sonnet 5, it is available with zero data retention.
Claude Opus 5.5
Anthropic states that Opus 5.5 "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5". Anthropic built it for long-running agentic coding and knowledge work. External evaluators, including Frontier Design and METR, tested it before release.
Opus 5.5 also costs less per token than Opus 5. Input now costs $4 per million tokens against $5 on Opus 5, and output costs $20 against $25. Cache reads fall to $0.20 from $0.50 and cache writes fall to $5 from $6.25, while batch processing runs at half the standard price, at $2 for input and $10 for output.
Anthropic adds that Opus 5.5 generates output more than 30% faster than Opus 5. A fast mode in Claude Code and on the Claude Platform runs at up to 2.5 times that speed. It costs $8 per million input tokens and $40 per million output tokens. For subscribers, Anthropic is raising five-hour usage limits on the Pro and Max plans and on Team and seat-based Enterprise plans. Subscribers also get a rate limit reset that they can save and use when they choose.
The announcement also describes a change in how the model communicates. Anthropic states that Opus 5.5 communicates more naturally than prior models and "puts the most important information up front", which answers feedback the company received about Opus 5. Its prompting guide adds that in knowledge work the model is much less likely to state an incorrect figure or cite the wrong source.
Anthropic's benchmark table compares Opus 5.5 with Opus 5 and with Fable 5.1.
Terminal-Bench 4.0. 66.4% for Opus 5.5, 52.3% for Opus 5 and 55.8% for Fable 5.1.
FrontierCode v1.1, Main. 54.4% for Opus 5.5, 48.0% for Opus 5 and 50.3% for Fable 5.1.
CursorBench 4.0. 57.8% for Opus 5.5, 46.6% for Opus 5 and 51.8% for Fable 5.1.
GDPval-AA v2.1. Scores of 1846 for Opus 5.5, 1708 for Opus 5 and 1735 for Fable 5.1.
OSWorld 2.1, partial. 81.8% for Opus 5.5, 74.0% for Opus 5 and 80.7% for Fable 5.1.
The announcement qualifies these figures, stating that "benchmark margins have become a less reliable guide to real-world differences". It adds that in Anthropic's own use, "the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest". Like Fable 5.1, Opus 5.5 carries Anthropic's watermarking measures for the EU AI Act, and this analysis of the Claude watermark looks at those measures in detail.
Sonnet 5.5 and Opus 5.5 Side by Side
The two models have the same context window and output limit, and the same knowledge cutoff. Their differences lie in price, speed, thinking behaviour and the work each one targets.
Price per million tokens. Sonnet 5.5 costs $2 for input and $10 for output, against $4 and $20 for Opus 5.5.
Relative speed. Anthropic's models overview rates Sonnet 5.5 as fast and Opus 5.5 as moderate.
Default effort. On the API, Sonnet 5.5 defaults to
highand Opus 5.5 tomedium, and in Claude Code both default tomedium.Thinking. Both models use adaptive thinking, where the model decides how much to reason. Developers can reduce Sonnet 5.5 to its lowest setting,
between_tools, while Opus 5.5 always thinks.Coding benchmarks. Opus 5.5 leads on FrontierCode, 54.4% against 46.2%, and on CursorBench, 57.8% against 55.5%. Sonnet 5.5 leads on Terminal-Bench 4.0, at 70.6% against 66.4%.
Knowledge work. On GDPval-AA, Sonnet 5.5 scores 1844 and Opus 5.5 scores 1846.
Computer use. On OSWorld 2.1, Sonnet 5.5 scores 80.1% and Opus 5.5 scores 81.8%.
Anthropic's Sonnet 5.5 prompting guide states that "for the hardest long-horizon work, an Opus model is the better choice". For API users who are unsure, the models overview recommends starting with Opus 5.5 for most workloads. It suggests Fable 5.1 when Opus 5.5 at higher effort still falls short. The AI model reference guide gives a wider view for readers weighing Claude against open-weight alternatives.
Using the New Models in Claude Code and the Claude Apps
According to the Claude Code documentation, Opus 5.5 is now the default model in Claude Code on paid plans and on the Anthropic API. Sonnet 5.5 needs Claude Code version 2.1.284 or later, and Opus 5.5 needs version 2.1.280 or later. In the terminal, the claude update command installs the latest version.
In the VS Code extension, Sonnet 5.5 appears greyed out in the model menu until you update the extension. The video below shows the process. Open the Extensions panel in VS Code, find Claude Code and select the update prompt on the extension. After it installs, restart extensions when VS Code asks. Sonnet 5.5 then appears in the model menu at the bottom of the chat window, and typing /model opens the same list.
Effort is the setting that controls how long the model reasons before it answers. Anthropic's Sonnet 5.5 announcement states that Claude Code and the Claude apps default to Medium effort. The Claude Platform defaults to High. The Claude Code documentation gives medium as the default for both new models. According to the same announcement, lower settings give faster answers with fewer tokens, which suits routine work, and higher settings let Claude reason for longer and check its work more thoroughly. In Claude Code, the /effort command opens a slider with five levels from low to max, and the level you pick there becomes the default for later sessions.
At xhigh and max, Anthropic's Sonnet 5.5 prompting guide notes, the model can start its own review rounds after a task. Those rounds sometimes use subagents and take more time and tokens, so the guide recommends high or below for routine work. The guide to reducing Claude Code token usage covers how effort and other settings affect token use.
Both models also tie their thinking to the account that produced it. Anthropic added this as a safeguard against distillation, where attackers use large numbers of fake accounts to extract a model's capabilities. Anthropic notes that most developers will not notice the change. Anyone who switches accounts partway through a Claude Code session should read its documentation on the change.
What Changes for API Integrations
Developers who call the models through the Claude API or a cloud provider face several breaking changes. The model IDs are claude-sonnet-5-5 and claude-opus-5-5, with an anthropic. prefix on Amazon Bedrock. Anthropic's "What's new" pages for Sonnet 5.5 and Opus 5.5 list the changes below.
Turning thinking off. A request that sends
thinking: {"type": "disabled"}now returns a 400 error on both models. Sonnet 5.5 acceptsbetween_toolsin its place athigheffort or below. Opus 5.5 has no way to turn thinking off.Forced tool use. Setting
tool_choicetoanyor to a named tool returns a 400 error on both models. For schema-valid tool input, Anthropic recommendsautowith strict tool use, or structured outputs.Thinking blocks. The API ties each thinking block to the model and the conversation that produced it. For accounts opened on or after 31 August 2026, editing earlier history and then replaying a thinking block returns a 400 error. Anthropic advises keeping conversations append-only.
Computer use. On the Claude API and Google Cloud, both models reject the earlier
computer_20251124tool and need thecomputer_toolset_20260801toolset, although Amazon Bedrock still accepts the earlier tool.Advisor tool. A Sonnet 5.5 executor rejects Opus 4.8, Opus 4.7 and Sonnet 5 as advisors.
Progress text. Notes the model writes between tool calls now arrive in
thinkingblocks, which are empty by default. An interface that streams those notes goes quiet until it sets adisplayvalue.Effort. Anthropic has recalibrated effort on both models. The Opus 5.5 API default drops to
mediumfrom thehighdefault on Opus 5. Anthropic advises re-running effort tests before carrying settings over.
Both models support prompt caching from a 512-token minimum prompt, and both work with batch processing and the Files API. Anthropic commits to keeping Opus 5.5 available on its own platforms until at least 22 September 2027. The same commitment runs to 28 September 2027 for Sonnet 5.5. The migration guides for Sonnet 5.5 and Opus 5.5 show the code before and after each change.
Safeguards and Refusals
Both models run safety classifiers that can decline a request. On the API, a declined request returns a normal response with a stop reason of refusal. The response also carries a category naming the policy area. Anthropic's documentation lists five categories, covering cyber harm, biological harm, help with building competing AI models, attempts to extract the model's internal reasoning, and general harms.
For cybersecurity, Anthropic states that both models still help users find and fix bugs in their own code. Higher-risk cybersecurity tasks on Sonnet 5.5 "visibly fall back to Sonnet 5". Opus 5.5 re-routes most cybersecurity tasks to Opus 4.8.
For biology, Opus 5.5 uses the same safeguards as Fable 5.1, and Sonnet 5.5 keeps the set used on Sonnet 5. Anthropic notes that most research, education and clinical work is unaffected on Sonnet 5.5. Some microbiology and virology requests may still trigger the classifier.
The Claude Code documentation describes what happens after a decline inside Claude Code. From Opus 5.5, cybersecurity requests re-run on Opus 4.8 and biology requests re-run on Opus 5. From Sonnet 5.5, cybersecurity requests re-run on Sonnet 5. A biology decline on Sonnet 5.5 ends with a refusal, because Sonnet 5.5 has no biology fallback model.
Organisations whose legitimate work meets these safeguards can apply to two Anthropic programmes. The Life Sciences Verification Program covers research and development in biology. The Cyber Verification Program gives cyberdefenders tiered access to more advanced capabilities, and Anthropic says it will soon extend to both new models.
Limitations
Hard Limits
Thinking. Opus 5.5 cannot run with thinking switched off, and Sonnet 5.5 can only reduce it to
between_tools.Forced tool use. Neither model accepts a
tool_choicethat forces a tool call through the API.Knowledge cutoff. Both models have a reliable knowledge cutoff of June 2026, so later facts need a search or a source.
Declined requests. Classifier declines apply in the five categories above. On the API, server-side fallback does not retry reasoning-extraction declines on either model.
Documented Weaknesses of Sonnet 5.5
Anthropic's prompting guide for Sonnet 5.5 documents the following behaviours.
Answers from memory. On chat and knowledge work, it sometimes answers from training knowledge when a search would catch changed details, and the guide's examples concern what is allowed or required and what is charged.
Stopping early. At
lowandmediumeffort, it sometimes pauses mid-task to confirm a plan or ask whether to continue.Unchecked changes. At
loweffort, it sometimes reports a code change as done without running a check. One example in the guide is skipping the tests because the project's dependencies are missing.Unrequested additions. It tends to add tests, documentation and small supporting files that nobody asked for. This happens at every effort level, and more often at higher effort.
Building too early. On an open-ended request, it can start building a presentation or report when the user only wanted ideas.
Mid-task messages. It sometimes treats a genuine message the user sends mid-task as a possible prompt injection. It may then ignore the message or ask the user to confirm it.
Dense visuals. It misses detail in dense charts and technical drawings unless it has a tool to crop or zoom the image.
Documented Weaknesses of Opus 5.5
Anthropic's prompting guide for Opus 5.5 documents a different set of behaviours.
Early stops in unattended runs. On long tasks, some of its progress updates end the turn. An automated agent loop that reads that as the end of the task stops there.
Missing context. In workflows across several connected apps, it tends to start work quickly. It can then miss information held somewhere the task did not name, such as an old email thread.
Revisiting earlier answers. In multi-turn chat, it sometimes goes back over an earlier answer while thinking about a new message, which adds latency.
Generic design. Without design direction, its frontend work falls back on a few default styles.
Anthropic's alignment assessment adds a broader caveat about Opus 5.5. The announcement reports signs that the model often suspects it is under evaluation. Anthropic says this makes it harder to assess how Opus 5.5 will behave in the settings where people use it.
Summary
Claude Sonnet 5.5 and Claude Opus 5.5 share a one-million-token context window and a June 2026 knowledge cutoff. Sonnet 5.5 keeps the Sonnet 5 price of $2 and $10 per million tokens and runs over 30% faster. Opus 5.5 drops to $4 and $20 and, by Anthropic's account, performs at the level of Fable 5.1 on most work. Anthropic positions Sonnet 5.5 for well-scoped everyday work and Opus 5.5 for complex work that needs careful judgement.
In Claude Code, both models default to medium effort, and Sonnet 5.5 appears once Claude Code is on version 2.1.284 or later. Developers moving existing code need to handle the breaking changes listed above before they switch model IDs. Registration for the GenAI Skills Academy community, where members discuss updates like these, is free at learn.genaiskills.io.
Sources
All sources verified 29 September 2026.
Introducing Claude Sonnet 5.5. Release date, positioning, pricing, speed, benchmarks, effort defaults and safeguards for Sonnet 5.5
Introducing Claude Opus 5.5. Release date, pricing, fast mode, usage limits, benchmarks, safeguards and the evaluation caveat for Opus 5.5
Models overview. Lineup, relative speed, context window, output limit, knowledge cutoff, default effort and retirement dates
What's new in Claude Sonnet 5.5. Breaking changes, feature support and refusal handling for the new Sonnet 5.5 model
What's new in Claude Opus 5.5. Breaking changes, batch pricing and behaviour differences for the new Opus 5.5 model
Prompting Claude Sonnet 5.5. Effort guidance and the documented weaknesses of Sonnet 5.5 in coding and knowledge work
Prompting Claude Opus 5.5. Effort guidance and the documented weaknesses of Opus 5.5 in unattended runs and multi-app workflows
Claude Code model configuration. Default model, version requirements, effort defaults and classifier fallback in Claude Code