Efficiency is real work.
Fireworks says Ember-1 matches Kimi K3 quality while using 40 percent fewer tokens. Its team reports more than 50 training experiments, over 200 evaluations, external benchmarks, live customer tests, and internal coding workloads.[1] That is a useful target. Shorter reasoning can cut cost and latency when the result still holds.
The claim comes from the model vendor. Its post describes the evaluations and results, but this garage did not reproduce the model comparison. Treat the number as a reason to test Ember-1 on your workload, not as a transferable discount.
A cheaper draft is cheaper. It is not more finished.
Even the word "Done" needs wiring.
Visual Studio Code's 1.140 Insiders notes include a small fix with a large lesson. Marking an active agent session as Done now stops that session so it does not keep consuming tokens in the background.[2] Before the fix, the label and the running process could disagree.
The same notes add a command that sends one prompt to multiple agents in isolated worktrees. A judge agent can recommend a result. Recommendation is useful routing. The human still needs a release boundary that names the chosen patch, checks, runtime result, and owner.
More drafts can make review worse.
Alex Ewerlöf argues that cheaper code creation does not remove maintenance, reliability, security, scalability, or accountability. He also calls out parallel agent output that grows faster than a person can review it.[3] This is an opinion essay, not a benchmark. Its strongest point is operational. Generation volume and review capacity are different measurements.
A team can improve model efficiency and still increase total waste if extra drafts enlarge the comparison set, hide unreviewed changes, or reach production without an owner. The useful denominator is the accepted change that survives its stated conditions.
Put four brakes on the release lane.
- Behavior. Name the requirement and attach the exact check that exercises it.
- Operations. Record rollout, stop, rollback, and the signal that triggers each action.
- Maintenance. Identify the code path, dependency, and future change that will make this patch expensive.
- Ownership. Name the person who accepts the result and the evidence they reviewed.
Optimize tokens inside that boundary. Do not use token savings to erase it.