Pimp My IDE / test deck
Back to garage
October 4, 2026 | Agent testing / replay / assertions

Record the route. Test the arrival.

TesterArmy's e2e lets an agent find a path through an interface. A later exact assertion can qualify that path for replay without another model call.

The useful split is simple. Let the model handle uncertain navigation. Let a separate check decide whether the product did the right thing.

The model gets the uncertain part.

The e2e package mixes natural-language agent steps with ordinary locators and assertions. In an agent.act step, the model receives a redacted text snapshot of screen roles, names, text, and state. The runner authorizes each action, applies a budget, records the result, and asks for a verdict before the deadline.[1]

That makes the agent useful when the route is easier to describe than to script. It can find the upgrade path even if the current menu arrangement differs from the last build. The runner can also stop repeated failures and request a verdict when the budget runs low.[1]

Use the agent to find the route. Do not ask the route to certify itself.

The witness comes after the action.

The cache records an agent.act sequence only after a later check verifies the result. The documentation counts locator assertions, engine assertions, waits, agent.assert, and agent.waitFor as verification. A plain value comparison, a read, another action, or the attempt merely finishing does not qualify the recording.[2]

That distinction matters. "Click whatever upgrades the workspace" is an instruction. "The status reads Pro" is an observation. One chooses actions. The other can fail those actions.

When the expected value is exact, the project's assertion guide recommends a locator assertion. It retries without a model call. A model-backed assertion is for a judgment that cannot be reduced to one exact value, and inconclusive evidence has its own failure code.[3]

Replay needs a clutch.

On a later run, e2e can find recorded controls by role, name, test id, and nearby context. It repeats the actions, then checks the final route and the state changes seen when the recording was made. If a target is missing, ambiguous, or no longer produces the expected end state, replay stops and the agent can take over from the current screen.[2]

This is better than treating a recorded script as permanent truth. It also creates a new review job. Named controls, clean test data, exact postconditions, and a visible no-cache run become part of the test contract.

Cached does not mean fresh.

A replayed action path proves that the recorded route still matched its checks. It does not prove that the model could rediscover the path today. The CLI exposes --no-cache for a fresh agent run. Keep at least one scheduled or pre-release fresh run if discoverability matters to the product.

The package is real and installable. The npm registry reported e2e@0.17.0 on October 4. We ran its help command in an isolated temporary home. The CLI printed commands for initialization, test runs, exploration, replay-cache inspection, MCP access, and telemetry control. We did not configure a provider, launch a browser, or run a product test.[4]

Telemetry is on by default. The project documents the fields and supports E2E_TELEMETRY_DISABLED=1 and DO_NOT_TRACK=1. Set the policy in CI instead of leaving every runner to decide.[5]

Interactive makeover / replay witness rack

Give the tape a witness head.

This replaces a vague "AI test passed" note with one mode selector, four review interlocks, a causal route, and a copyable contract. It prepares a test plan. It does not run the framework or verify your app.

Transport controls

Plan builder

Pick the run that this card requests. Then select only the interlocks that your test definition includes. Selection does not mean the test passed.

Requested run mode
Test-definition interlocks

Witness transport

Structure only
TARGETOPEN
ACTOPEN
WITNESSOPEN
FRESHNESSOPEN
0/4
Test definition open

No interlock is selected. The requested mode is record.

This rack drafts a contract. It does not exercise your app, qualify a replay entry, inspect cache files, contact a model, or prove a postcondition.

Sources read

Source log and evidence boundary
  1. TesterArmy e2e, "How agent steps work", read October 4, 2026. It documents the redacted screen snapshot, runner and executor split, action budgets, loop stops, verdicts, and secret handling.
  2. TesterArmy e2e, "Caching agent steps", read October 4, 2026. It documents verification-qualified recording, replay matching, hand-off behavior, cache modes, exact failure reasons, and the no-cache run.
  3. TesterArmy e2e, "Assertions", read October 4, 2026. It separates exact locator assertions from model-backed judgments and documents inconclusive evidence.
  4. npm registry metadata for e2e 0.17.0 and the source repository, read October 4, 2026. The registry supplied version 0.17.0 and its executable. We ran npx --yes e2e@0.17.0 --help in an isolated temporary home.
  5. TesterArmy e2e, "Telemetry", read October 4, 2026. It documents default-on usage telemetry, the listed fields, debug inspection, and opt-out controls.
  6. Playwright, "Locators", read October 4, 2026. It recommends user-facing attributes and explicit contracts such as roles, and notes that locators resolve the current DOM element before each action.

Evidence boundary. TesterArmy documents the framework behavior. npm supplied the published CLI artifact. Playwright documents locator behavior used by the web engine. Pimp My IDE designed the witness rack and review method. This page did not run a browser test, verify a replay cache entry, or compare models.