A product failure exposed the right problem.
Ayman Nadeem built Nuanced around persistent plans for coding agents. In his September 24 post, he says early users had little appetite for long generated specs. He also says the split between planning and building made the work feel staged. Implementation kept revealing questions that the finished plan had declared settled.[1]
That is one builder's report about one product. It does not prove that every plan mode has failed. It does name a familiar failure. The plan becomes a document to approve instead of a record that changes when the system answers back.
Keep the decision that can still change. Drop the prose that only proves a planning step happened.
The vendors still keep planning inside the loop.
GitHub's current cloud-agent documentation says the agent can research a repository, create an implementation plan, change a branch, run tests, accept follow-up questions, and iterate before a pull request.[2] Planning is present, but it sits beside execution and review.
Anthropic's agent guide describes agents as models that use tools, read environmental feedback, and pause at checkpoints or blockers. It recommends fixed workflows for predictable tasks and flexible agents when the required steps cannot be known in advance.[3]
OpenAI's desktop documentation puts parallel work, files, browser actions, diffs, and follow-up chat in one workspace. The page tells users to create and inspect outputs, then refine them.[4] That is a work loop, not a promise that the first outline will survive contact with the repository.
Replace the plan handoff with a decision handoff.
A long plan tries to predict the whole route. A next-decision card names only the current move and the evidence needed before the next move. It stays short because its job is narrow.
- Frame. Name the user-visible result, the boundary, and the person who can change either one.
- Probe. Make the smallest reversible change or inspection that can disprove the current idea.
- Inspect. Read the diff, test output, runtime behavior, and new unknowns.
- Decide. Continue, revise, stop, or ask for judgment. Record why.
The loop can run in minutes for a small fix or across several sessions for a migration. Its size follows uncertainty, not task prestige.
Human attention needs a target.
Parallel agents make chat history a poor control panel. The operator should not read every internal thought. The operator should see the decisions that change scope, cost, access, architecture, or user behavior.
The ratchet below keeps one active phase, one question, one evidence request, and a short lap tape. It does not inspect a repository or verify a build. It produces a card for the next real tool run.