Pimp My IDE / garage dispatch
Back to garage
September 28, 2026 | Sonnet 5.5 / effort / output budgets

Save fuel for the answer.

Claude Sonnet 5.5 makes effort cheaper and faster by Anthropic's measurements. The harder lesson is in the controls. Thinking, tool calls, and final text share one output ceiling.

A bigger effort setting can spend more of the same tank. Define the output ceiling, final-answer reserve, checkpoint, and stop rule before the run starts.

The new engine is real. The numbers are vendor numbers.

Anthropic released Claude Sonnet 5.5 on September 28. The company says it is more than 30 percent faster than Sonnet 5 and can cost up to 30 percent less per task in its tests. The listed API price remains $2 per million input tokens and $10 per million output tokens. Anthropic also reports a 70.6 percent Terminal-Bench 4.0 score, compared with 10.3 percent for Sonnet 5.[1]

Those figures describe Anthropic's evaluation setup. They do not predict the cost, success rate, or review burden for your repository. GitHub made Sonnet 5.5 generally available in Copilot on the same day, so the migration question already reaches ordinary editor workflows.[4]

Faster reasoning still needs enough runway to land an answer.

Effort draws from every output channel.

Anthropic's effort documentation says the setting affects every output token. That includes visible text, tool calls, function arguments, and thinking. Sonnet 5.5 supports low, medium, high, xhigh, and max. High is the documented default. Max permits the model to spend without an effort constraint, but the request still has a hard max_tokens ceiling.[2]

The model overview lists a 128,000-token maximum output. That is different from the one-million-token context window. Context is the room available to read the conversation. Output is the shared tank for thinking and the response.[3]

One public run hit the redline before delivery.

In the Hacker News launch thread, Simon Willison posted one effort sweep for an SVG prompt. His max-effort run consumed 128,000 thinking tokens over 15 minutes and 40 seconds, then ended without the requested final SVG. His low through xhigh runs returned responses. This is one public observation, not a controlled benchmark or a general failure rate.[5]

The observation matters because it demonstrates a concrete failure shape. A run can spend its output allowance on internal work and leave no deliverable. The operator needs a deadline, a retry ceiling, and a fallback gear. A larger token ceiling alone does not create that policy.

A model upgrade can change the wiring.

Anthropic's Sonnet 5.5 migration guide says thinking is on by default. It replaces the earlier disabled mode with between_tools for the lowest thinking setting. The guide also warns that max_tokens covers thinking plus text and that thinking tokens are billed as output. Code that assumes the first content block is text can break because a response may begin with a thinking block.[6]

The same guide changes forced tool use. Sonnet 5.5 rejects tool_choice values that force any tool or one named tool. It calls for automatic selection and strict schemas where supported. A model ID swap can alter control flow, parsing, and cost at once.

Install an output reserve.

  1. Set effort from task difficulty and evidence, not from release-day excitement.
  2. Set a hard output ceiling and wall-clock deadline.
  3. Reserve space for the requested answer in your operating policy. The API does not enforce a separate final-answer reserve.
  4. Checkpoint long tool loops before they consume the whole run.
  5. On timeout, budget pressure, or repeated tools, shift down or stop. Do not restart at max by reflex.
  6. Replay a fixed workload across effort levels. Measure completed work, total output, elapsed time, tool calls, and review burden.
Interactive makeover / reasoning fuel reserve

Keep runway for delivery

Traditional purpose replaced: one effort dropdown with no operating boundary. Better version: choose the effort gear, declare an output ceiling and final-answer reserve, close four run-policy breakers, and copy the exact test card.

Set the run policy

These controls write a test plan. They do not call a model or predict token use.

Effort gear
32,000 tokens
4,000 tokens128,000 tokens
25%
10% policy floor60% policy floor
Run-policy breakers
Policy draft1 of 4 breakers selected
Operator-declared budget

Reasoning fuel reserve

Planned final reserve8,000
Remaining shared allowance24,000

The run policy needs three more sections.

One deadline section is selected. The reserve is a planning target, not an API partition.

What this component proves. It calculates a user-declared planning reserve inside a selected output ceiling and writes a test-card structure. It does not reserve provider tokens, measure a model, estimate quality, enforce a stop, or attach run evidence.

Sources and limits

Open the source log
  1. Anthropic, "Introducing Claude Sonnet 5.5", September 28, 2026. Read September 28. This first-party launch supplies the speed, price, per-task cost, benchmark, capability, and safety claims. It also says benchmark scores capture only one part of model capability.
  2. Anthropic Claude Platform, effort documentation, read September 28, 2026. It defines the effort levels, says effort affects all output tokens, and documents task guidance, defaults, and cost controls.
  3. Anthropic Claude Platform, model overview, read September 28, 2026. It lists the Sonnet 5.5 model ID, price, adaptive thinking, high default effort, one-million-token context window, and 128,000-token maximum output.
  4. GitHub Changelog, "Claude Sonnet 5.5 in GitHub Copilot", September 28, 2026. GitHub says the model is generally available in Copilot. Availability does not validate Anthropic's benchmark or cost claims.
  5. Hacker News launch discussion for Sonnet 5.5, September 28, 2026. Simon Willison's comment supplies one public max-effort observation and a comparison across effort settings. It is an anecdotal run, not a controlled evaluation.
  6. Anthropic Claude Platform, Sonnet 5.5 migration guide, read September 28, 2026. It documents model ID, thinking-mode, output-budget, content-block, tool-choice, and schema changes.

Evidence boundary. Anthropic's release metrics are first-party reports. GitHub verifies Copilot availability. The Hacker News observation covers one prompt and one operator. The interactive reserve is a planning template. Its percentage does not create a separate provider-side token pool.