The new engine is real. The numbers are vendor numbers.
Anthropic released Claude Sonnet 5.5 on September 28. The company says it is more than 30 percent faster than Sonnet 5 and can cost up to 30 percent less per task in its tests. The listed API price remains $2 per million input tokens and $10 per million output tokens. Anthropic also reports a 70.6 percent Terminal-Bench 4.0 score, compared with 10.3 percent for Sonnet 5.[1]
Those figures describe Anthropic's evaluation setup. They do not predict the cost, success rate, or review burden for your repository. GitHub made Sonnet 5.5 generally available in Copilot on the same day, so the migration question already reaches ordinary editor workflows.[4]
Faster reasoning still needs enough runway to land an answer.
Effort draws from every output channel.
Anthropic's effort documentation says the setting affects every output token. That includes visible text, tool calls, function arguments, and thinking. Sonnet 5.5 supports low, medium, high, xhigh, and max. High is the documented default. Max permits the model to spend without an effort constraint, but the request still has a hard max_tokens ceiling.[2]
The model overview lists a 128,000-token maximum output. That is different from the one-million-token context window. Context is the room available to read the conversation. Output is the shared tank for thinking and the response.[3]
One public run hit the redline before delivery.
In the Hacker News launch thread, Simon Willison posted one effort sweep for an SVG prompt. His max-effort run consumed 128,000 thinking tokens over 15 minutes and 40 seconds, then ended without the requested final SVG. His low through xhigh runs returned responses. This is one public observation, not a controlled benchmark or a general failure rate.[5]
The observation matters because it demonstrates a concrete failure shape. A run can spend its output allowance on internal work and leave no deliverable. The operator needs a deadline, a retry ceiling, and a fallback gear. A larger token ceiling alone does not create that policy.
A model upgrade can change the wiring.
Anthropic's Sonnet 5.5 migration guide says thinking is on by default. It replaces the earlier disabled mode with between_tools for the lowest thinking setting. The guide also warns that max_tokens covers thinking plus text and that thinking tokens are billed as output. Code that assumes the first content block is text can break because a response may begin with a thinking block.[6]
The same guide changes forced tool use. Sonnet 5.5 rejects tool_choice values that force any tool or one named tool. It calls for automatic selection and strict schemas where supported. A model ID swap can alter control flow, parsing, and cost at once.
Install an output reserve.
- Set effort from task difficulty and evidence, not from release-day excitement.
- Set a hard output ceiling and wall-clock deadline.
- Reserve space for the requested answer in your operating policy. The API does not enforce a separate final-answer reserve.
- Checkpoint long tool loops before they consume the whole run.
- On timeout, budget pressure, or repeated tools, shift down or stop. Do not restart at max by reflex.
- Replay a fixed workload across effort levels. Measure completed work, total output, elapsed time, tool calls, and review burden.