Pimp My IDE / garage dispatch
Back to garage
September 28, 2026 | agents / output cost / release proof

Cheap output still needs a finish line.

A faster model can lower the cost of a draft. It cannot decide when the code is ready to own.

Count tokens during generation. Count behavior, operations, maintenance, and ownership before release.

Efficiency is real work.

Fireworks says Ember-1 matches Kimi K3 quality while using 40 percent fewer tokens. Its team reports more than 50 training experiments, over 200 evaluations, external benchmarks, live customer tests, and internal coding workloads.[1] That is a useful target. Shorter reasoning can cut cost and latency when the result still holds.

The claim comes from the model vendor. Its post describes the evaluations and results, but this garage did not reproduce the model comparison. Treat the number as a reason to test Ember-1 on your workload, not as a transferable discount.

A cheaper draft is cheaper. It is not more finished.

Even the word "Done" needs wiring.

Visual Studio Code's 1.140 Insiders notes include a small fix with a large lesson. Marking an active agent session as Done now stops that session so it does not keep consuming tokens in the background.[2] Before the fix, the label and the running process could disagree.

The same notes add a command that sends one prompt to multiple agents in isolated worktrees. A judge agent can recommend a result. Recommendation is useful routing. The human still needs a release boundary that names the chosen patch, checks, runtime result, and owner.

More drafts can make review worse.

Alex Ewerlöf argues that cheaper code creation does not remove maintenance, reliability, security, scalability, or accountability. He also calls out parallel agent output that grows faster than a person can review it.[3] This is an opinion essay, not a benchmark. Its strongest point is operational. Generation volume and review capacity are different measurements.

A team can improve model efficiency and still increase total waste if extra drafts enlarge the comparison set, hide unreviewed changes, or reach production without an owner. The useful denominator is the accepted change that survives its stated conditions.

Put four brakes on the release lane.

  1. Behavior. Name the requirement and attach the exact check that exercises it.
  2. Operations. Record rollout, stop, rollback, and the signal that triggers each action.
  3. Maintenance. Identify the code path, dependency, and future change that will make this patch expensive.
  4. Ownership. Name the person who accepts the result and the evidence they reviewed.

Optimize tokens inside that boundary. Do not use token savings to erase it.

Interactive makeover / hydraulic release brake

Done line brake test

Traditional purpose replaced: one Done checkbox. Better version: choose the boundary, route four evidence lines into one brake, and copy the missing fields. This bench builds a review template. It does not test a patch.

Set the stopping boundary

The native radio and checkbox controls own the state. The hydraulic display mirrors the selected review structure.

Boundary under review
Brake open / 0 of 4 lines selected

Teaching proxy. Pressure shows selected template sections, not test coverage.

Sections to include in the handoff
No evidence section is selectedThe candidate boundary is set. Choose the fields this handoff must request.
Copyable done line card

Record the stop, proof, and owner

Replace every required field with evidence from your own run.

What this component proves. It keeps generation state, stop status, release proof, and ownership in one handoff template. It does not inspect workers, run tests, compare patches, or approve a release.

Sources and limits

Open the source log
  1. Fireworks AI, "Introducing Ember-1", published September 23, 2026 and read September 28, 2026. The vendor reports its token-reduction claim, training process, benchmark results, customer tests, and internal use.
  2. Visual Studio Code 1.140 Insiders release notes, updated September 24, 2026 and read September 28, 2026. Microsoft lists the Done-session stop fix, external-session labels, and multiple-agent comparison command.
  3. Alex Ewerlöf, "Coding is NOT solved", published September 26, 2026 and read September 28, 2026. The article is an attributed opinion about code generation, review volume, nonfunctional requirements, and accountability.
  4. Hacker News discussion for "Coding is NOT solved", item 49877988, read September 28, 2026. This was a discovery trail and reader discussion. It does not verify the article's claims.

Evidence boundary. We did not benchmark Ember-1, test the current VS Code Insiders build, or measure review outcomes. The brake test is a planning aid. Its selected sections are requests for evidence, not evidence that work passed.