Sonnet 5.5 keeps Sonnet 5's $2/$10 pricing, runs 30% faster and nears Opus 5.5 on benchmarks. How to set effort, cache and batch to cut cost per task.
Claude Sonnet 5.5 replaces Claude Sonnet 5 at exactly the same price: $2 per million input tokens, $10 per million output tokens, $0.20 per million cached reads. Same tokenizer, same 1M-token context window, same 128K output ceiling. What changes is how much work each of those tokens buys.
Same price, more work per dollar
Anthropic released Sonnet 5.5 on September 28, 2026, six days after Opus 5.5. Its pitch is cost per task, not cost per token: the model generates output more than 30% faster than Sonnet 5 and, Anthropic says, costs up to 30% less per task because it needs far fewer tokens and model requests to finish the same work.
The benchmark gap over Sonnet 5 is wide, and on several tests Sonnet 5.5 lands within a few points of Opus 5.5 at half the price:
| Benchmark (Anthropic-reported) | Sonnet 5 | Sonnet 5.5 | Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 10.3% | 70.6% | 66.4% |
| CursorBench 4.0 | 34.1% | 55.5% | 57.8% |
| FrontierCode 1.1 (main) | 42.4% | 46.2% | 54.4% |
| OSWorld 2.1 (computer use) | 57.0% | 80.1% | 81.8% |
| Humanity's Last Exam (with tools) | 54.9% | 64.5% | 67.7% |
| GDPval-AA v2.1 (knowledge work) | 1449 | 1844 | 1846 |
| Chartography (no tools) | 15.6% | 61.6% | 64.4% |
The efficiency claims are sharper still. At low or medium effort, Anthropic says Sonnet 5.5 already beats Sonnet 5's best scores on several benchmarks at about one-tenth of the per-task cost. Its developer migration guide adds that on most agentic coding evals, Sonnet 5.5 at medium outscored Sonnet 5 at high at under a fifth of the cost, and on computer use it completed substantially more tasks than Sonnet 5 at its highest effort with under a third of the tokens.
These are vendor numbers; independent testing is still catching up, and Anthropic notes its Sonnet 5.5 scores came from a pre-release build with a since-fixed structured-output bug. But the direction is clear. With the rate card frozen and each job needing fewer tokens and fewer round-trips, the bill falls without any price cut.
It also sits at half the price of Claude Opus 5.5 ($4 / $20). For most everyday coding, agent and enterprise work, Sonnet 5.5 is now the default value pick; Opus remains the better choice for the hardest long-horizon tasks.
| Model | Input $/1M | Output $/1M | Cached read $/1M |
|---|---|---|---|
| Claude Sonnet 5.5 | $2.00 | $10.00 | $0.20 |
| Claude Sonnet 5 | $2.00 | $10.00 | $0.20 |
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 |
| Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 |
The most efficient way to run it
Six levers for running Sonnet 5.5 cheaply, in order of impact.
1. Set effort explicitly, and start low
Effort is the biggest lever, and the levels were recalibrated on Sonnet 5.5: medium no longer means what it meant on Sonnet 5. Do not carry old settings across. The API default is still high, so leaving it unset quietly overspends on simple traffic.
lowfor chat, content generation, classification, extraction and search. Atlowthe model skips thinking on most simple requests.mediumfor agentic coding and multistep tool use.xhighandmaxonly where you have measured a quality gain.
Be wary of max in particular. On FrontierCode, Sonnet 5.5 scored worse at max than at xhigh, because the extra effort triggered review loops that timed out. Simon Willison saw the same pattern in a simple SVG test: max burned all 128,000 output tokens ($1.28) and produced nothing, while xhigh finished in 41 seconds for under 6 cents. Note too that Claude Code now defaults Sonnet 5.5 to medium, while the API defaults to high.
client.messages.create(
model="claude-sonnet-5-5",
max_tokens=16000,
output_config={"effort": "low"},
messages=[{"role": "user", "content": "..."}],
)
From medium upward the model thinks briefly before nearly every reply, even a greeting. Asking it in the system prompt to "think less" has almost no effect; lowering effort does.
2. Cache everything that repeats
Cached reads cost $0.20 per million, a tenth of the normal input price. The minimum cacheable prompt drops to 512 tokens on Sonnet 5.5, down from 1,024 on Sonnet 5, so shorter system prompts and tool lists now qualify.
Keep stable content first (system prompt, tool definitions) and volatile content last. A timestamp or request ID in the system prompt silently breaks the cache on every call. Check usage.cache_read_input_tokens: if it stays at zero across repeated calls, something in the prefix is changing.
3. Change effort without breaking the cache
Changing the top-level effort between requests invalidates the prompt cache. Sonnet 5.5 supports per-message effort (beta mid-conversation-output-config-2026-07-01), which keeps it. Run an interactive session at low and raise effort only for the one hard turn.
Sonnet 5.5 also accepts mid-conversation system messages, which Sonnet 5 did not. Append new operator instructions to the message list instead of editing the top-level system prompt, and the cached history stays valid.
4. Batch anything that can wait
Nightly reports, bulk classification, backfills: the Message Batches API runs at half price, and Sonnet 5.5 keeps Sonnet 5's batch rates.
5. Use Opus as an advisor, not as the worker
The advisor tool lets a Sonnet 5.5 executor consult a stronger model only when it needs to. Pair it with Opus 5.5, Opus 5, or Fable 5/5.1. Opus 4.8, Opus 4.7 and Sonnet 5 advisors are now rejected. You pay Opus rates for the occasional consult rather than for every token.
6. Judge cost per finished task
A cheaper request that needs three retries is not cheaper. Log requests-per-task and tokens-per-task, not just price per call. That is the metric on which Sonnet 5.5 separates from Sonnet 5.
What breaks when you switch
The price is the same; the request surface is not. Five changes will return a 400 on code written for Sonnet 5:
thinking: {type: "disabled"}is rejected. Try adaptive thinking atlowfirst. If a route must stay thinking-off, sendthinking: {type: "between_tools"}at efforthighor below.- Forced tool use is rejected.
tool_choiceanyandtoolreturn a 400. Useauto, mark the toolstrict: true, and steer from the prompt — or use structured outputs when you only wanted JSON back. - Thinking blocks are bound to the model and the conversation. No other model reads Sonnet 5.5's thinking blocks, and editing earlier turns invalidates them. Keep your message history append-only.
- Computer use needs
computer_toolset_20260801on the Claude API and Google Cloud. - The advisor tool accepts fewer advisors (see above).
Existing Sonnet 5 prompts otherwise work unchanged. Also remove old workarounds — refusal steering, tool-call retry shims, "do not be lazy" — and re-run your evals; Anthropic says the model declines fewer benign requests and uses tools more reliably.
The security angle
Sonnet 5.5 declines in five stop_details categories: cyber, bio, frontier_llm, reasoning_extraction and general_harms. A decline is an HTTP 200 with stop_reason: "refusal", so always check stop_reason before reading content.
Finding vulnerabilities in source code is allowed; malware and exploit development trigger the cyber classifier. Anthropic says higher-risk cybersecurity tasks visibly fall back to Sonnet 5. Server-side fallback (fallbacks: "default", beta server-side-fallback-2026-07-01, Claude API only) retries cyber and frontier_llm declines on Sonnet 5. Teams doing legitimate offensive security work should look at Anthropic's Cyber Verification Program.
Sonnet 5.5 also runs in its own rate-limit pool, separate from Sonnet 5. Check your tier's limits before moving volume, and note that Priority Tier is not supported.
Bottom line
Swap claude-sonnet-5 for claude-sonnet-5-5, fix the five breaking changes, set effort explicitly (low for chat, medium for agents), cache aggressively and batch what can wait. Same rate card, fewer tokens per finished job.