Pimp My IDE / Garage Dispatch
Back to garage
September 24, 2026 | agent skills / held-out gates / decision records

Stop promoting prompt edits on vibes.

An agent procedure is executable policy. Treat each rewrite like a release candidate. Score it, limit the change, test it on work it has not seen, and keep the rejected version in the log.

The take: The next useful IDE panel will not be another prompt box. It will be a test stand where a procedure must beat its baseline without erasing the rules that already work.
Load the Skill Ratchet Test Stand

The instruction file is part of the machine.

Teams edit agent instructions as if they were notes taped above a monitor. Add a warning after a failure. Rewrite a section when a new model lands. Remove a constraint because the file feels long. The edit may help one task while damaging another.

SkillOpt treats the skill document as trainable state for a fixed agent. A separate optimizer proposes bounded add, delete, and replace operations. The paper's default path accepts a candidate only when it strictly improves a held-out validation score.[1] That mechanism matters more than the benchmark headline. A procedure change has to earn adoption.

Small edits need a hard gate.

The SkillOpt authors use a textual learning-rate budget to limit edit size. They also retain rejected edits as feedback and use slower updates across epochs.[2] In plain terms, do not let one bad run trigger a rewrite of the whole operating manual.

A correction is a candidate. It becomes policy after a matched test, not after a persuasive explanation.

The reported results cover six benchmarks, seven target models, and three execution harnesses. The paper says SkillOpt was best or tied across its 52 evaluated cells.[1] Those are paper results on named tasks and setups. They do not prove that automated skill editing will improve an unrelated repository, model, or team.

Human understanding belongs in the gate.

Whiteboard takes a different route to the same pressure point. Its repository describes a canvas where an agent can connect diagrams and trace quotes to code, an AST-aware diff view, and a decision log for autonomous choices.[3] It also names current limits, including weak multi-repository review and shared reviews that do not update after publication.

A score can catch task regression. It cannot tell a maintainer whether the new rule remains legible, whether its exception is understandable, or whether the team can defend the decision later. Put the candidate skill beside its diff and decision record. Make a person inspect the policy they are about to inherit.

Borrow deployment discipline.

devenv 2.4 applies a related pattern to machine changes. Its release post describes saved plans, stale-plan rejection when the target or system generation changes, confirmation before apply, and target-side rollback for failed activation or health checks.[4]

An agent skill needs the same separation. Build the candidate. Review the diff. Test the exact candidate against a pinned baseline and held-out fixtures. Adopt the tested file without silently rebuilding it. Keep the previous version ready for rollback.

Run the promotion drill.

  1. Pin the target model, harness version, tools, permissions, budget, seed policy, and baseline skill hash.
  2. Collect scored successes and failures. Do not train on anecdotes without an oracle.
  3. Limit the candidate to named add, delete, or replace operations. Record the diff.
  4. Evaluate baseline and candidate on the same held-out fixtures. Keep those fixtures out of the edit loop.
  5. Require both score improvement and human review of the procedure diff.
  6. Adopt the exact tested candidate. Save the rejection record and rollback hash.
Interactive makeover / procedure release gate

Skill Ratchet Test Stand.

Traditional purpose replaced: edit an instruction file and trust the next run. Better version: lock four promotion requirements, generate a test card, and keep the final state at TEST SPEC READY until real scores and hashes are attached.

Set the promotion pawls

Each circuit closes one requirement. The carriage shows specification progress, not measured agent quality.

Promotion circuits
DRAFT0 / 4 circuits
0requirements specified

The rewrite has no release gate.

Start by pinning scored rollouts and the baseline skill hash.

Measured quantityrequirements selected
Candidate qualitynot measured
Adoption authorityhuman review
Why it is better: the score source, edit limit, unseen test, and rollback record stay separate. A green stand means the release test has been specified. It does not mean the candidate passed.

Print the promotion card

Sources read, not vibes

Open the source log
  1. Yang et al., "SkillOpt: Executive Strategy for Self-Evolving Agent Skills": frozen target agent, bounded text edits, strict held-out validation acceptance, evaluated task matrix, and reported results.
  2. SkillOpt project page: rollout, reflection, edit, and gate loop; textual learning-rate budget; rejected-edit buffer; slow update; ablation tables; and transfer summary.
  3. Whiteboard repository: local open-source review canvas, code-linked diagrams, AST-aware diff viewer, decision log, privacy notes, and current limitations.
  4. devenv 2.4 Machines release post: reviewable deployment plans, stale-plan rejection, confirmation, target-side activation lock, health checks, rollback scope, and stated recovery limits.
  5. Hacker News discussion 49833867: the current Whiteboard launch thread and practitioner discussion. It does not verify the repository's product claims.

Source boundary: SkillOpt's gains belong to the paper's models, benchmarks, harnesses, and evaluation method. Whiteboard describes its own product behavior. devenv documents machine deployment behavior, not agent-skill validation. This dispatch combines their mechanisms into an operating proposal. The test stand is a teaching tool, not production telemetry.