The mid-tier is not fallback-free

Same list price is not a silent model swap.

On September 28, 2026, Anthropic introduced Claude Sonnet 5.5 as the second model in the Claude 5.5 family after Opus 5.5. The launch post prices it at $2 per million input tokens and $10 per million output tokens, matching Sonnet 5, with $0.20 per million for cache reads. Anthropic says Sonnet 5.5 typically needs far fewer tokens for the same work (up to 30% less cost per task in its testing) and generates outputs 30%+ faster than Sonnet 5. It positions Sonnet 5.5 as a faster, lower-cost complement to Opus 5.5: Opus for complex work that needs careful judgment, Sonnet for well-scoped everyday tasks, bug fixes, and polished documents, slides, and spreadsheets (Anthropic, September 28, 2026).

That framing is the easy half. The harder half is that Sonnet 5.5 is the first Sonnet model to launch with cybersecurity safeguards and fallbacks like those Anthropic built for its most capable models. Higher-risk cyber requests can visibly fall back to Sonnet 5 mid-conversation. Biology and distillation refusals can block without a fallback. API customers do not get automatic switching by default. Developers who ran Sonnet 5 with thinking off must move to the new between_tools setting before the ID swap will even accept the request. Default effort is high on the Claude API and medium in Claude Code (Anthropic; Claude Platform docs; Claude Help Center; claude.dev, all accessed or published September 28 to 29, 2026).

This is a primary-source review, not a hands-on bakeoff. The operator job is a migration checklist, a fallback and billing expectation, and a refusal-posture rewrite. Do not treat “same price” as same bill or same mid-tier behavior.

What it is

Sonnet 5.5 is Anthropic’s speed-and-intelligence mid-tier in the 5.5 family. Model ID on the Claude API and most clouds is claude-sonnet-5-5 (Bedrock uses anthropic.claude-sonnet-5-5). Context window is 1M tokens. Max output is 128K tokens on the synchronous Messages API, with a higher batch beta ceiling documented separately. Knowledge cutoff is June 2026. Adaptive thinking is on by default. Effort levels run low through max. Platform docs put default effort at high on the Claude API; Claude Code defaults Sonnet work to medium (Claude Platform Sonnet 5.5 overview; claude.dev, September 28, 2026).

Anthropic’s launch table and narrative put Sonnet 5.5 far above Sonnet 5 on agentic coding and knowledge-work scores, and close to Opus 5.5 on some evaluations, while still saying Opus remains clearly stronger on complex, open-ended judgment work. Early tester quotes (Epic, Slack, Zendesk, Unity, and others in the launch post) emphasize fewer steps, fewer tokens, and faster iteration. Treat those as vendor-selected signals, not your eval. Haiku 5.5 is promised for high-volume work in the coming weeks, so the family still has a cheaper tier arriving (Anthropic, September 28, 2026).

What changed

Efficiency at unchanged list price. The headline price did not move. Anthropic’s claim is that token use per task dropped enough to cut cost up to about 30% in its testing, with 30%+ faster output generation. claude.dev restates the same price table versus Opus 5.5 and warns that surfaces ship different default efforts, so your bill moves with effort and token shape even when the rate card looks identical (Anthropic; claude.dev, September 28, 2026).

Thinking defaults and between_tools. On Sonnet 5.5, omitting the thinking field runs adaptive thinking. Sending thinking: {"type": "disabled"} returns a 400. The replacement for “thinking off” is between_tools, which keeps up-front thinking off and only thinks between tool calls. It works at low, medium, and high effort. At xhigh or max it errors. It rejects companion fields like display or budget_tokens. Platform docs and the migration guide treat this as a breaking change for anyone who disabled thinking on Sonnet 5 (Claude Platform overview and migration guide; claude.dev).

Five other migration breaks. Forced tool choice of type any or tool errors; use auto plus strict tools. Thinking blocks are tied to model and conversation. On the Claude API and Google Cloud, the older computer_20251124 computer-use tool is rejected in favor of the newer toolset form. The advisor tool rejects several older advisor pairings. Separately, text between tool calls can arrive in thinking blocks, so UIs that assumed content[0].text go quiet until they read blocks by type (Claude Platform overview; claude.dev).

Cyber safeguards land on the mid-tier. Because Anthropic rates Sonnet 5.5’s cyber capabilities as a large step up from Sonnet 5 (comparable in the launch narrative to Opus-class cyber capability), Sonnet 5.5 ships with cyber classifiers and fallbacks. Help Center: higher-risk offensive cyber requests (exploit generation, binary-based vulnerability scanning, penetration testing) may fall back to Sonnet 5. Source-code vulnerability scanning for secure coding is described as still allowed on Sonnet 5.5. Biology safeguards match Sonnet 5 and block without fallback. Frontier LLM development classifiers can fall back to Sonnet 5. Distillation and reasoning-extraction attempts block outright. Classifiers review not only the latest user message but memory, connector content, web search results, and files, so content nobody typed can trigger a switch (Claude Help Center; Anthropic launch post; The New Stack, September 28, 2026).

Apps auto-switch. API does not (by default). In Claude apps, automatic model switching is on by default. After a cyber or frontier-LLM fallback, the chat continues on Sonnet 5 with a visible notice, and the picker stays on Sonnet 5 until you change it. You can disable automatic switching in settings; then a flagged request pauses instead of switching. On the Claude API, automatic switching is not active by default. Without configured fallbacks, a declined request can return HTTP 200 with stop_reason: "refusal" and category detail rather than silently continuing on Sonnet 5. Server-side fallback (fallbacks: "default", beta) retries cyber and frontier_llm declines on Sonnet 5; it does not retry bio, reasoning_extraction, or general_harms (Claude Help Center; claude.dev; The New Stack, September 28, 2026).

Isometric ice-blue chip at a fork with one path diverting in soft mint

How it works (and breaks)

Fallback is a product behavior, not a bug. If your agent loop assumes every Sonnet 5.5 request stays on Sonnet 5.5 for the whole session, you will mis-attribute quality, latency, and cost when a classifier diverts. Help Center is explicit that after an automatic switch, returning to Sonnet 5.5 can fall over again if the triggering content is still in the thread. Editing the prior message before retry is the documented mitigation.

Billing across a fallback is not “same price, one line.” Help Center: blocks before any output may or may not be charged depending on classifier category (biology, distillation, and LLM development refusals before output can be billed; some other pre-output blocks are not). Midstream blocks charge input and streamed tokens at the producing model’s rates. If automatic switching sends the chat to Sonnet 5, that Sonnet 5 response is charged separately at Sonnet 5 rates, with a credit described for the cache miss of the fallback request. API teams that enable server-side fallback need to read the refusal and fallback docs before they promise finance a flat Sonnet 5.5 run rate (Claude Help Center; claude.dev refusals section).

Effort defaults will surprise teams that only swap the model ID. API default high versus Claude Code default medium means the same prompt can think longer and cost more on one surface than the other. claude.dev advises re-running effort sweeps because levels are recalibrated versus Sonnet 5. If you are pushing xhigh or max on Sonnet for hard jobs, Anthropic’s own guidance is to consider whether Opus 5.5 is the better SKU rather than forcing the mid-tier into Opus-shaped work.

Secure coding is not a free pass for every security workflow. Anthropic draws a line between source-code vulnerability work (allowed) and higher-risk offensive patterns (fallback or block). Binary-focused scanning and exploit generation sit on the wrong side of that line in the Help Center examples. Agent systems that pull advisories, binaries, or web content into context can trip classifiers on retrieved text, not just on the user prompt (Help Center; The New Stack).

Who it is for

Good fit: teams whose daily load is well-scoped coding, iterative product work, document and slide polish, and repeated agent tasks with a clear check for done. Anthropic and claude.dev both steer that work to Sonnet 5.5 and reserve Opus 5.5 for harder judgment and long-horizon ambiguity.

Poor fit without extra design: security research workflows that look like offensive cyber to a classifier; biology-adjacent work that already hit Sonnet 5 false positives; API products that disabled thinking and have not migrated to between_tools; any finance model that assumed a pure ID swap at identical unit economics with no fallback lines.

Not a hands-on verdict: this review does not claim we reproduced Terminal-Bench scores, cyber fallback rates, or token deltas on production traffic. Vendor evals and early tester quotes are directional. Your migration gate should be your own golden tasks plus refusal and fallback telemetry.

Two blank geometric chips with a graphite barrier and diverting mint path

Pricing and limits that matter this week

List price parity with Sonnet 5 is real on the rate card: $2 / $10 input / output per MTok, cache read $0.20, with matching batch and cache-write structure called out in the launch and builder posts. Token shape is not parity. Fewer tokens per task is the whole cost story. Faster tokens change latency budgets and streaming UX. High default effort on the API can erase “cheaper mid-tier” if you never lower effort on routine jobs. Claude Code’s medium default is a different starting point again (Anthropic; claude.dev; Platform overview).

Limits that trip migrations: disabled thinking is invalid; between_tools cannot ride at xhigh/max; forced tool_choice shapes error; older computer-use tool forms error on Claude API and Google Cloud; advisor pairings can error; temperature and related sampling knobs at non-default values return 400 on this model per Platform docs. Plan the checklist before you flip a production alias.

What to do this week

  1. Inventory thinking-off callers. Every client that sent `thinking: {“type”: “disabled”}` on Sonnet 5 needs `between_tools` (and a legal effort level) before it can speak to Sonnet 5.5. Add a content-block-by-type reader so thinking progress updates do not break UIs.
  1. Decide fallback policy per surface. Apps: leave auto-switch on for end users, or turn it off if you need hard stops and manual model choice. API: decide whether to opt into server-side default fallback, SDK middleware, or explicit refusal handling. Document that cyber and frontier_llm can retry on Sonnet 5, while bio and distillation do not.
  1. Add fallback fields to FinOps. Track model ID on the response, refusal category, whether a fallback credit applied, and separate Sonnet 5.5 versus Sonnet 5 spend. “Same $2/$10” is not a forecast if fallovers and effort shifts move token volume.
  1. Re-run effort on golden tasks. Start from API high or Code medium as documented defaults, then sweep. Do not assume Sonnet 5’s old effort setting means the same work. For hard open-ended jobs, evaluate Opus 5.5 instead of cranking Sonnet to max.
  1. Rewrite the security narrative. Tell security and platform teams Sonnet 5.5 can divert mid-conversation on higher-risk cyber content, including content pulled in by tools. Secure coding against source remains in-bounds per Anthropic; binary and exploit-shaped work may not stay on 5.5. Watch for Cyber Verification Program expansion if you need tiered access later.
  1. Ship the migration as a checklist, not a banner. Model ID swap, thinking migration, tool_choice, computer-use toolset, advisor pairing, refusal handling, effort sweep, FinOps tags. Anything less is a silent ID swap fantasy.

Sonnet 5.5 is a real mid-tier upgrade on Anthropic’s numbers: faster, thriftier tokens, closer to Opus on some scoped work, still priced like Sonnet 5 on the rate card. It is also the first Sonnet that can fall over like an Opus-class safety stack. The mid-tier is not fallback-free. Migrate the parameters, meter the divert, and do not brief “same price, drop-in” to anyone who owns uptime or the bill.