Pimp My IDE / garage dispatch
Back to garage
September 28, 2026 | faster models / review capacity / proof

The model got faster. The check lane did not.

Claude Sonnet 5.5 arrives with vendor claims of higher speed and lower task cost. Good. Put the saved time into checking the work instead of filling the merge queue faster.

Route every release claim to one fitting check: read the source, recompute the number, run the behavior, or name the person making the judgment.

Generation speed moves the bottleneck.

Anthropic says Claude Sonnet 5.5 produces output more than 30 percent faster than Sonnet 5 and costs up to 30 percent less per task in its tests. The list price remains $2 per million input tokens and $10 per million output tokens. Anthropic attributes the lower task cost to using fewer tokens for the same work.[1]

Those are vendor measurements. They still matter. A faster model can shorten the draft loop. It can also increase the number of changes that reach review in one afternoon. If review capacity stays fixed, generation speed turns into a longer queue.

Spend the speed gain on evidence, not output volume.

One disclaimer cannot check four kinds of claims.

Glyph's essay on serious AI products makes a useful product argument. A warning that AI can make mistakes gives the user a job without giving them a workbench. The essay proposes visible claim checks, readable citations, human notes, and clear data provenance.[2]

The check depends on the claim. A quotation needs the original source. A number needs a repeatable calculation. A behavior claim needs an executable test. A design or policy decision needs a named owner and stated tradeoff. A generic "reviewed" badge hides these differences.

Automated review is another producer.

GitHub's current Copilot code review documentation separates model credits from the runner time used for repository context and tools. It also says reviews still run in a more limited form when the supporting Actions path fails. The same page prices Lite and Balanced review at different estimated ranges.[3]

That is a useful reminder. Review output has a mode, a cost, an execution path, and a failure state. Save those facts. Do not treat an automated comment as independent proof merely because it arrived in the review column.

Make the check visible before merge.

  1. Copy the exact claim from the generated patch, summary, or review.
  2. Classify it as source, number, behavior, or judgment.
  3. Attach the matching evidence. Do not substitute a test for a product decision or a citation for runtime behavior.
  4. Record who checked it and what remains open.
  5. Keep the claim open until the evidence can be inspected by the next reviewer.

The release should get faster only when the check lane gets faster too.

Interactive makeover / check lane pit board

Put each claim on the right lift

Traditional purpose replaced: one vague "verify this" instruction. Better version: classify the claim, close the matching evidence clamps, and copy a review card that keeps missing proof visible.

Route the claim

The claim type changes the evidence request. The four native checkboxes record which sections belong in the review card.

Claim type
Evidence clamps
Copyable check card

Draft the evidence route

Replace every bracketed field. Selecting every clamp prepares the card structure. It does not prove the claim.

What this component proves. It builds one review-card structure from the claim type and selected evidence sections. It does not fetch a source, recompute a number, run a test, appoint a reviewer, or approve a release.

Sources and limits

Open the source log
  1. Anthropic, "Introducing Claude Sonnet 5.5", September 28, 2026. Read September 28. This first-party launch post supplies the speed, pricing, task-cost, benchmark, and tester claims. Its benchmark table and cost plots are vendor evaluations.
  2. Glyph, "What Would A Serious AI Product Look Like?", September 27, 2026. Read September 28. This is an opinion essay. It argues for first-class mistake checking, visible citations, human notes, structured task interfaces, and data provenance.
  3. GitHub Docs, "About GitHub Copilot code review", read September 28, 2026. The page documents review surfaces, Lite and Balanced effort, estimated credit ranges, Actions runner use, and the limited fallback when supporting workflows fail.
  4. Claude Platform Docs, "Effort", read September 28, 2026. The page documents effort as a control over response token use and lists different defaults and task fits across current models.
  5. Hacker News discussion for Claude Sonnet 5.5, September 28, 2026. Used for discovery and public reaction only. Comments do not verify model behavior.

Evidence boundary. This article did not run Claude Sonnet 5.5 or compare it with another model. Speed, cost, and benchmark figures belong to Anthropic's launch evaluation. The pit board is a review template. It does not inspect evidence or measure a production workflow.