Pimp My IDE / garage dispatch
Back to garage
October 5, 2026 | agents / orchestration / run evidence

Your workflow is code. Its run is still evidence.

GitHub's new dynamic workflows move multi-agent control flow out of a one-off prompt and into a program. That makes the route reviewable. It does not make every branch deterministic, every pause durable, or every result proved.

Version the route. Limit the run. Save the receipt.

Control flow has joined the repository.

GitHub released dynamic workflows in public preview for Copilot CLI, the Copilot app, and the Copilot SDK. A workflow can combine commands, tools, services, user checkpoints, and one or more agents. The author defines the steps, conditions, and handoffs in code. Agents handle work that needs analysis or judgment.[1][2]

That split matters. A coded branch can guarantee that two reviews are requested before a merge step. It cannot guarantee that either reviewer understood the change. The program controls when judgment runs and what shape it must return. The model still supplies the judgment.

Repeatable routing is not repeatable reasoning.

Pick the control owner on purpose.

GitHub distinguishes dynamic workflows from autopilot and /fleet. Autopilot keeps working without asking after each step. Fleet lets Copilot decide how to divide parallel work. A dynamic workflow follows steps and handoffs defined by its author.[2]

OpenAI's Agents SDK makes the same design choice explicit. An LLM can choose tools and handoffs for an open task. Code-driven orchestration can instead classify structured output, chain agents, run independent work in parallel, or loop an evaluator around another agent. OpenAI describes the code-driven route as more predictable for speed, cost, and performance.[3]

Use model judgment inside a bounded step. Use code for the order, branch condition, retry count, spend stop, approval point, and final record.

A budget field is not always a hard stop.

GitHub documents limits for concurrent workflow-owned agents, total agents launched, active running time, and approximate AI credits. It also says credit usage arrives after consumption, so work already running can take the total past the configured maximum.[2]

Write that distinction into the run card. Concurrency is a queue limit. Total agents and time can stop a run. The credit value is an approximate stop, not a billing ceiling. Put the provider's real account-level cap beside it if spend must not cross a number.

Pause and replay are different promises.

GitHub says a paused or limit-stopped workflow can retain status and saved results. A resumed run can reuse those saved results. Work that was not saved may run again. Sharing a workflow shares its definition, not the original run history or saved progress.[2]

Temporal shows the stronger durability contract that teams often assume without naming it. Its workflow engine records an ordered event history and replays code to rebuild state. External calls run as activities whose recorded results are reused during replay.[4] GitHub does not claim that model for dynamic workflows. If your process must survive a crash without repeating a payment, deployment, or message, define the idempotency and replay contract yourself.

Interactive makeover / workflow flight recorder

Choose who steers. Arm four records.

Traditional purpose replaced: an orchestration dropdown plus a separate checklist. Better version: native controls drive one visible mode deck, four independent witness lines, and a copyable review card. No agent run starts here.

Set the control mode

The labels compare planning ownership. They do not rank model quality or predict task success.

Steering mode
Flight records to request
0 of 4 records selectedCoded workflow
Mode deck / witness bank

Workflow flight recorder

Control ownerAuthor + code
Path shapeDefined
Run proofMissing

The mode is selected. No run record is selected.

Next record: definition pinned.

All four selected records means the review-card structure is ready. It does not prove that the workflow ran, stayed inside a limit, paused safely, preserved state, or produced a correct result.

Definition / judgment / stop / receipt

Four things to put under review.

01 / DEFINITION

Version the route

Keep stages, branches, schemas, prompts, tool access, and model names beside the code revision.

02 / JUDGMENT

Mark model decisions

Name every step where an agent classifies, writes, approves, rejects, or chooses the next route.

03 / STOP

Separate each limit

Record queue width, total launches, active time, provider spending, and the person who may resume.

04 / RECEIPT

Save what happened

Keep the run ID, inputs, outputs, approvals, errors, external side effects, and final revision.

Sources read

Source log and evidence boundary
  1. GitHub Changelog, "Dynamic workflows in Copilot CLI and the Copilot app", published October 1 and read October 5, 2026. It announces the public preview, supported surfaces, code-defined process model, example uses, and CLI setup.
  2. GitHub Docs, "Dynamic workflows", read October 5, 2026. It documents planning ownership, permissions, sharing, revisions, run limits, approximate credit behavior, and pause or resume semantics.
  3. OpenAI Agents SDK, "Agent orchestration", read October 5, 2026. It separates LLM-led orchestration from code-led routing and documents structured outputs, chaining, evaluator loops, and parallel execution.
  4. Temporal Docs, "Temporal Workflow", read October 5, 2026. It explains workflow definitions, executions, ordered event history, deterministic replay, and activity-result reuse. We use it as an adjacent durability reference, not as a description of GitHub's implementation.

Evidence boundary: GitHub documents a public-preview product and its controls. OpenAI documents orchestration patterns in its SDK. Temporal documents its own durable execution model. Pimp My IDE designed the recorder and review card. We did not install Copilot CLI, author an extension, start a dynamic workflow, measure credits, interrupt a run, resume one, or compare output quality.