Pimp My IDE / Garage logBack to dispatches
Agent safety04 Oct 20266 min read
Garage log 193 / computer-use execution

Batch action fusebox

A model can return several clicks and keystrokes in one turn. That does not make the sequence one atomic move. Your runner owns order, failure, skipped work, and the next screenshot.

The practical call: treat every action batch as a committed prefix. Run in order. Stop at the first error. Record the skipped suffix. Inspect the state that the successful prefix left behind.
What changed

One response can contain a small program

Computer-use APIs can package several dependent actions together. The integration, not the model, executes those actions against the live interface.

Anthropic's current computer-use toolset can return several member tool calls in one model turn. Its documentation calls this a batch action. The application must run the calls in order because a later action can depend on the screen state created by an earlier one.

OpenAI's computer tool uses a similar shape. One computer_call can contain an ordered actions array. OpenAI tells the application to execute permitted actions in order, capture the updated screen, and return that observation. A call marked completed means the model finished generating it. The interface has not moved until the application executes it.

A batch is a route, not a transaction.

Failure leaves a real prefix

Suppose a batch asks the runner to focus a field, type a value, and take a screenshot. Focus succeeds. Typing fails because the window changes. The runner cannot pretend that nothing happened. The field may still hold focus, a menu may have opened, or a previous value may have been selected.

Anthropic's contract is explicit. Stop at the first failure. Return a normal result for each successful action. Return an error for the failed action. Return an error for every later action to say it was not executed. Every requested block still needs an answer.

The observation closes the loop

The next screenshot is not decoration. It is the evidence needed to plan from the state that now exists. OpenAI recommends a screenshot after a short group of actions. Anthropic notes that batches often end with one, and allows the application to attach one when they do not.

This matters because desktop state is unstable. GitHub's current computer-use documentation warns that timing and window changes can produce different results, repeat an action, or stop progress. GitHub also recommends a direct API, terminal command, filesystem tool, MCP server, or dedicated browser tool when one exists. Structured routes reduce the amount of state that pixels have to carry.

01 / ORDER

Run the array as written

Do not parallelize actions that share focus, pointer position, or window state.

02 / TRIP

Stop at the first error

Do not let an action planned for the old state run against the new one.

03 / ANSWER

Return every result

Mark the successful prefix, failed action, and unexecuted suffix separately.

04 / INSPECT

Read the state again

Capture the screen or use a stronger independent check before continuing.

Interactive makeover / sequence fusebox

Trip the action bus

A normal activity spinner hides partial execution. This teaching rig shows which actions passed, which action failed, and which later actions never ran.

Choose an injected result

This rig does not control a desktop or call a model. It simulates result bookkeeping for one fixed action list.

Physical state / ordered execution rail

Focus, type, inspect

The receipt is a simulation record. It does not prove a production runner implements this contract.

Injected failure / typeNot run
01FOCUSPending
02TYPEPending
03SCREENSHOTPending
04RESULT SETPending
Simulation has not run.

Choose a fault, then run the fixed action list.

Shop notes

Put the fuse in the runner

Prompt wording can describe the policy. Only the execution code can enforce it when the interface changes between actions.

Validate each action before dispatch. Apply confirmation rules before each consequential block, not once for the whole batch. A batch can cross a consent boundary within one turn.

Keep the action result and the state check separate. A successful click means input was delivered. It does not prove the intended record changed. Use an application readback, file check, API query, or human review when the consequence matters.

Then test the ugly prefixes. Fail action one, the middle action, and the final observation. Confirm that the runner never executes the suffix, always answers every requested action, and resumes from a fresh view. That is where the batch contract becomes more than a diagram.

Sources read

Open the receipts

The product contracts come from first-party documentation. The fusebox and committed-prefix language are Pimp My IDE's implementation model.

Evidence boundary: these sources describe three separate products and integration contracts. They do not prove equivalent implementations or reliability. No computer-use model or desktop runner was exercised for this dispatch. The interactive fusebox only simulates a four-state bookkeeping rule.