The instruction file is part of the machine.
Teams edit agent instructions as if they were notes taped above a monitor. Add a warning after a failure. Rewrite a section when a new model lands. Remove a constraint because the file feels long. The edit may help one task while damaging another.
SkillOpt treats the skill document as trainable state for a fixed agent. A separate optimizer proposes bounded add, delete, and replace operations. The paper's default path accepts a candidate only when it strictly improves a held-out validation score.[1] That mechanism matters more than the benchmark headline. A procedure change has to earn adoption.
Small edits need a hard gate.
The SkillOpt authors use a textual learning-rate budget to limit edit size. They also retain rejected edits as feedback and use slower updates across epochs.[2] In plain terms, do not let one bad run trigger a rewrite of the whole operating manual.
A correction is a candidate. It becomes policy after a matched test, not after a persuasive explanation.
The reported results cover six benchmarks, seven target models, and three execution harnesses. The paper says SkillOpt was best or tied across its 52 evaluated cells.[1] Those are paper results on named tasks and setups. They do not prove that automated skill editing will improve an unrelated repository, model, or team.
Human understanding belongs in the gate.
Whiteboard takes a different route to the same pressure point. Its repository describes a canvas where an agent can connect diagrams and trace quotes to code, an AST-aware diff view, and a decision log for autonomous choices.[3] It also names current limits, including weak multi-repository review and shared reviews that do not update after publication.
A score can catch task regression. It cannot tell a maintainer whether the new rule remains legible, whether its exception is understandable, or whether the team can defend the decision later. Put the candidate skill beside its diff and decision record. Make a person inspect the policy they are about to inherit.
Borrow deployment discipline.
devenv 2.4 applies a related pattern to machine changes. Its release post describes saved plans, stale-plan rejection when the target or system generation changes, confirmation before apply, and target-side rollback for failed activation or health checks.[4]
An agent skill needs the same separation. Build the candidate. Review the diff. Test the exact candidate against a pinned baseline and held-out fixtures. Adopt the tested file without silently rebuilding it. Keep the previous version ready for rollback.
Run the promotion drill.
- Pin the target model, harness version, tools, permissions, budget, seed policy, and baseline skill hash.
- Collect scored successes and failures. Do not train on anecdotes without an oracle.
- Limit the candidate to named add, delete, or replace operations. Record the diff.
- Evaluate baseline and candidate on the same held-out fixtures. Keep those fixtures out of the edit loop.
- Require both score improvement and human review of the procedure diff.
- Adopt the exact tested candidate. Save the rejection record and rollback hash.