The token price has not changed
Anthropic released Claude Sonnet 5.5 on September 28, 2026. It is available in Claude and through the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Foundry. The API model identifier is claude-sonnet-5-5.
Standard pricing remains $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 per million tokens, while five-minute cache writes cost $2.50 and one-hour cache writes $4 per million.
Batch API processing receives a 50% input and output discount. Anthropic lists a context window of up to one million tokens, 128,000 tokens of standard maximum output and a 300,000-token maximum for Batch API in beta.
Thirty percent cheaper is a per-task claim
That distinction matters when forecasting an API bill. Anthropic is not offering a 30% discount on each million tokens. It says Sonnet 5.5 generally consumes fewer tokens to perform the same work and can therefore cost up to 30% less per task than Sonnet 5.
The speed claim needs similar qualification. Anthropic says output generation is more than 30% faster. That is a vendor measurement rather than a guarantee that an entire application will finish every workflow 30% sooner. External tools, network latency, reasoning length and effort settings remain part of the execution path.
Early customers quoted by Anthropic report efficiency gains of the same general kind. Slack says its offline evaluations used about 14% fewer output tokens. Lovable reports roughly one-third fewer tool calls and about half as many shell runs. Box says its own workloads were 2.4 times faster while consuming 12% fewer total tokens.
Published scores put Sonnet unusually close to Opus
Anthropic reports 70.6% for Sonnet 5.5 on Terminal-Bench 4.0, an agentic coding evaluation, compared with 10.3% for Sonnet 5 in the same launch table. Opus 5.5 is listed at 66.4% using the Xhigh setting that produced its best result.
On CursorBench 4.0, Sonnet 5.5 reaches 55.5%, versus 34.1% for Sonnet 5 and 57.8% for Opus 5.5. GDPval-AA v2.1 gives the two newer models Elo scores of 1,844 and 1,846 respectively.
Those numbers do not establish that Sonnet replaces Opus. Anthropic itself says Opus 5.5 remains clearly stronger on complex, open-ended work requiring sustained judgment. Launch benchmarks cover specific tasks, and much of the performance evidence is selected or published as part of Anthropic's own release material.
One benchmark actually gets worse with more effort
FrontierCode 1.1 offers a useful counterexample to the assumption that spending more compute must improve a result. Anthropic reports 46.2% for Sonnet 5.5 at Max effort but 52.1% at Xhigh.
Anthropic attributes the reversal to model behavior at Max. Sonnet 5.5 invoked Claude Code's review skill more frequently, splitting work among multiple subagents. In two cases examined by Cognition, this produced either a timeout or additional edits beyond the requested scope, which FrontierCode penalizes.
That is a practical agentic limitation. More effort can mean more actions, and every additional action creates another opportunity to do something the task never requested.
Effort level is now an economic control
Anthropic exposes effort settings so developers and users can trade speed and token consumption against deeper reasoning. Claude Code and Claude apps default to Medium, while the Claude Platform defaults to High.
An application budget therefore cannot be derived from the $2/$10 rate alone. Two calls using the same model at the same token price can generate very different bills if one reasons longer, invokes more tools or produces substantially more output.
That is why the 30% saving has to be measured on the actual workload. An agent that avoids several searches, shell calls or correction loops may produce a meaningful reduction. A simple request that already required few tokens under Sonnet 5 has much less waste available to remove.
Enterprise numbers are impressive, but mostly private
Balyasny Asset Management says it evaluated Sonnet 5.5 on 2,441 private finance tasks spanning Q&A, extraction, analysis and forecasting. In the testimonial published by Anthropic, the new model used about 121,000 tokens per answer where Sonnet 5 used 497,000.
Base44 reports an average of 3.6 iterations across 118 real app builds, compared with 7.7 using Opus 5. Unity says Sonnet 5.5 completed 90% of tasks in its internal multi-step Unity Editor and coding benchmark.
These results are useful because they describe workflows closer to production applications than many isolated benchmark questions. They are still private evaluations summarized by the participating companies in Anthropic's launch material, and complete datasets and protocols are not public in every case.
A million-token context is not a million-token accuracy guarantee
Sonnet 5.5 accepts text and images and produces text. Anthropic specifies a context window of up to one million tokens and a reliable knowledge cutoff of June 2026.
Context size is not an accuracy metric. Feeding the model much larger collections of documents enables larger workflows, but it does not guarantee that every detail will be used correctly or that the resulting answer will be error-free.
For document-heavy systems, the less dramatic efficiency improvements may ultimately matter more: fewer steps and fewer tokens for comparable output. At high volume, eliminating a handful of unnecessary tool calls thousands of times is a real infrastructure change.
Sonnet now gets cyber safeguards previously associated with stronger models
Anthropic says Sonnet 5.5's cybersecurity capabilities have advanced enough to warrant protections similar to those deployed around Opus 5.5. Higher-risk cybersecurity requests can visibly fall back to Sonnet 5, while routine software development and ordinary bug fixing remain available.
The model keeps the biology safeguards used by Sonnet 5. Anthropic acknowledges that legitimate microbiology and virology requests can sometimes be flagged incorrectly and offers verification programs intended to provide approved organizations with appropriate access.
Anthropic also describes Sonnet 5.5 as the first Sonnet release equipped with classifiers intended to prevent industrial-scale distillation attacks that use large numbers of fake accounts to extract model reasoning capabilities.
Migration can require code changes even though pricing does not
Moving from Sonnet 5 to Sonnet 5.5 is not completely transparent for every integration. Anthropic says developers running Sonnet with thinking disabled need to move to the new between_tools setting before migrating if they want to keep upfront thinking switched off.
That detail captures the release rather well. The sticker price is unchanged, but model behavior, tool use, speed and safeguards have shifted enough to change both application economics and output.
At $2 input and $10 output, Sonnet 5.5 is half the per-token price of Opus 5.5 at $4 and $20. The practical question is therefore not whether the two models occasionally post similar benchmark scores. It is how many tasks in a real application genuinely require the additional judgment for which Anthropic still positions Opus.