Steel ruler resting diagonally on a scarred wooden workbench, its graduated tick marks sharp under warm overhead light, background in deep shadow

Sonnet 5.5’s Terminal-Bench Leap Is Real. The Score Has Two Problems.

/ Maxim Starkweather / 6 min read

Anthropic released Claude Sonnet 5.5 on Sunday, September 28. The headline number is Terminal-Bench 4.0: 70.6%, up from Sonnet 5’s 10.3%. That’s a 60-point jump, the largest single-generation improvement Anthropic has published on an agentic coding benchmark. It is also off by six points — and understanding why tells you more about what Anthropic built than the number itself. Five breaking API changes shipped alongside the model. Any production code targeting claude-sonnet-5 needs a migration pass before the first Monday agent run.

Two Numbers, One Benchmark

Terminal-Bench 4.0 measures a model’s ability to complete complex, multi-step professional tasks inside a command-line interface — the benchmark closest to what agentic coding tools actually do in production. Anthropic claims 70.6%. Artificial Analysis, which runs independent evaluations, measured 64%. The same gap appeared on Opus 5.5 — Anthropic’s number higher, Artificial Analysis’s lower — making this a pattern in the 5.5 series rather than a one-off calibration difference.

Artificial Analysis ran their evaluation on pre-release hardware and found a structured output bug in that deployment; they plan to rerun once the fix is in production. Their numbers may climb toward Anthropic’s. Even if the gap closes to two or three points, understanding what drove it is worth the time.

Within Anthropic’s own published benchmark table, Sonnet 5.5 scores above Opus 5.5 on Terminal-Bench — which is architecturally strange. Opus is Anthropic’s heavier, more capable model. The community thread on Hacker News surfaced the explanation quickly: in Terminal-Bench’s agentic trials, Opus 5.5 had approximately 10% of attempts routed to a fallback model because its safeguards declined the task. Sonnet 5.5, at the time of Anthropic’s evaluation, had roughly 1.5% fallback rate. In benchmark scoring, a fallback counts as a failure. Opus was losing points not because it lacked the capability, but because its safety layer was blocking tasks that Sonnet was allowed to attempt.

Artificial Analysis’s independent numbers, where both models sit at 60–64% range, tell the more likely story of how they compare when evaluated under consistent conditions. The net result: even at 64%, the jump from Sonnet 5’s 10.3% is a genuine 6x improvement on the hardest agentic coding benchmark in the current evaluation landscape. The improvement is real. The specific number in the headline is inflated — by pre-release testing conditions, a differing fallback regime, and a structured output bug that has since been fixed.

Terminal screen with an empty line where output should have appeared, representing the silent behavior change in Sonnet 5.5's thinking blocks

One more data point from Artificial Analysis that trade coverage has largely ignored: AA-Omniscience, their factual accuracy benchmark, shows Sonnet 5.5 at 54% versus Opus 5.5’s 66%. Sonnet 5.5 beats Opus at agentic terminal tasks. Opus beats Sonnet at factual recall by twelve points. The $4/$20-per-million-token price difference between Opus and Sonnet has a real capability tradeoff behind it, and it is not uniform across task types. Anyone building a knowledge-retrieval pipeline on Sonnet 5.5 because the Terminal-Bench number looked competitive with Opus is working from the wrong benchmark.

The Safety Architecture Explains the Score

Sonnet 5.5 is the first Sonnet model to ship with Opus-level cyber protections. That is the sentence that explains the fallback pattern, the Terminal-Bench gap, and most of what matters about this release.

Sonnet 5.5 ships with five named refusal categories accessible via stop_reason: "refusal" and a stop_details object. Two are new. "frontier_llm": the model declined because the request could assist the development of competing AI models. This is Anthropic’s anti-distillation measure, operating at the inference layer, not the policy layer. A prompt that tries to extract Sonnet 5.5’s weights or training behavior into a competing model hits this category. "reasoning_extraction": the request asked the model to reproduce its internal reasoning in the response text. A request asking Sonnet 5.5 to surface its chain-of-thought verbatim — as a way to understand or replicate its reasoning patterns — will hit this.

The existing categories remain: "cyber" for requests that could enable cyber harm, "bio" for biological harm, and "general_harms" for everything else under the usage policy. Higher-risk cybersecurity tasks — the ones Opus 5.5’s safeguards were already blocking at a 10% rate in Terminal-Bench — now trigger "cyber" and fall back to Sonnet 5 via server-side fallback, if you’ve enabled it.

Two hands in different-colored gloves reaching toward the same glowing object, separated by a hair of space — the gap between independent and vendor benchmark scores

The implementation detail that will catch builders off guard: a declined request returns HTTP 200. Not a 4xx. The response carries stop_reason: "refusal" and the stop_details object, but any application that checks only for a successful status code will process the refusal as a completed response. If your application streams the model’s output to users and checks for an error before rendering, a "cyber" refusal will render nothing, silently, with a success status. The exception: server-side fallback (fallbacks: "default", currently in beta on the Claude API) retries "cyber" and "frontier_llm" declines against Sonnet 5. It does not retry "bio", "reasoning_extraction", or "general_harms" declines. Those return HTTP 200 with a refusal and no retry, regardless of fallback configuration.

The token-use data from Artificial Analysis completes the picture. At maximum effort settings, Sonnet 5.5 used approximately 193,000 output tokens per task on their Intelligence Index evaluation — around 7x GPT-6 Astra’s maximum effort token use. Anthropic’s claim of 30% fewer tokens per task refers to production usage at default effort settings, not maximum effort benchmarking. Both claims are accurate; they describe different operating regimes. The implication for builders: if you’re running Sonnet 5.5 at effort: max for evaluations, your token cost is not the production cost you’d see at the default effort: high. Run your own effort sweep rather than carrying settings over from Sonnet 5.

Five Things That Break on Upgrade

None of the five breaking changes fail silently — they all return 400 errors pointing to the correct fix. One additional behavior change does fail silently, and it is the one most likely to affect user-facing applications before anyone notices.

The silent one first. Text the model writes between tool calls now comes back in thinking blocks, not text blocks. At the default display: "omitted", those blocks arrive empty. Any application that streams intermediate model reasoning to users — progress notes, status updates, the model explaining what it’s about to do — goes quiet between tool calls with no error and no indication that anything changed. The fix is to set a display value that returns the text, or to turn off up-front thinking with between_tools, which routes text back through text blocks instead.

The explicit 400s: sending thinking: {"type": "disabled"} returns a 400 pointing to between_tools. If your agent loop explicitly disables thinking to manage latency, that one-line change is required on upgrade. Forced tool use — tool_choice: {"type": "any"} or {"type": "tool", "name": "..."} — now returns a 400; the replacement is tool_choice: {"type": "auto"} with strict: true, or structured outputs for schema enforcement. On the Claude API and Google Cloud, the computer_20251124 computer use tool returns a 400; teams using computer use need to migrate to the computer_toolset_20260801 toolset. Amazon Bedrock still accepts the earlier tool. The advisor tool now requires a 5.5-era advisor when Sonnet 5.5 is the executor — pairings with Opus 4.8, Opus 4.7, or Sonnet 5 as advisors fail with a 400.

Two behavioral shifts that don’t throw errors: effort levels are recalibrated, so a medium or high effort setting on Sonnet 5.5 produces different thinking depth than the same setting on Sonnet 5. The guidance in the migration docs is to run a fresh effort sweep rather than carry prior settings. And thinking blocks are now account-bound — a block produced by Sonnet 5.5 on one account cannot be passed to a request from a different account. The API drops it silently and continues; the request succeeds without the accumulated reasoning. For systems that share thinking block histories across accounts or sessions — multi-tenant architectures where context is passed across organization boundaries — this is a non-obvious behavior change that will show up as unexpectedly degraded reasoning without any error.

The current model lineup now sits at Fable 5.1 ($10/$50 per million tokens), Opus 5.5 ($4/$20), Sonnet 5.5 ($2/$10), and Haiku 4.5 ($1/$5). Sonnet 5.5 and Opus 5.5 score within 2 Elo of each other on the GDPval-AA knowledge-work benchmark. On factual recall, Opus is 12 points ahead. On agentic terminal tasks, Sonnet is likely ahead — once Artificial Analysis reruns the Terminal-Bench evaluation on production hardware, the independent number will clarify how far ahead. The 6-point gap between Anthropic’s published score and AA’s independent measurement will almost certainly narrow. What won’t change is the architecture behind it: Anthropic built a significantly more capable agentic model and simultaneously tightened the boundary on what that capability can be applied to. The Terminal-Bench score and the refusal categories are two sides of the same release. Builders need both before they upgrade.

Steel ruler resting diagonally on a scarred wooden workbench, its graduated tick marks sharp under warm overhead light, background in deep shadow

AI-generated editorial illustration · TemperatureZero · September 29, 2026

Keep reading the signal

Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.

Subscribe Free

Continue the archive

Latest BriefingsArticlesAbout Temperature Zero