The useful idea is the handoff.
Oya says an agent can complete a browser task, save the run as a playbook, and replay it with new inputs without calling a model. Its site says the saved steps become Playwright code and free-text answers can still call a model when needed. It also says a changed page stops the replay at the affected step, lets an agent finish, and saves the repair for later runs.[1]
That is a better division of labor than paying a model to rediscover the same button path every night. The first run handles uncertainty. The replay handles repetition. The handoff between them deserves the same review as any generated integration.
"No model on replay" describes the engine. It does not prove the car reached the right address.
A locator is part of the contract.
Playwright's test generator records browser actions and produces code. Its documentation says the generator prefers role, text, and test ID locators. When a locator matches several elements, the generator tries to refine it to one target. The same tool can record visibility, text, and value assertions.[2]
Those choices matter after the demo. A coordinate says where a button was. A role and accessible name say what the button is. A result assertion says whether the click produced the state the job needed.
Waiting is behavior, not delay.
Playwright checks whether a target is visible, stable, able to receive events, and enabled before many actions. If those checks do not pass before the timeout, the action fails. Its assertions retry until the expected condition appears or the timeout expires.[3]
A fixed sleep can hide a race on one machine and waste time on another. A replay contract should name the condition that opens the next step. "Coverage active is visible" is useful. "Wait three seconds" is only a guess.
Authentication is cargo.
Playwright can save cookies, local storage, and IndexedDB state while recording. Its documentation warns that the saved file contains sensitive information and should stay local or be deleted after use.[2]
A replay system may manage identity differently, but the review questions stay concrete. Which profile runs the task? Which hosts may receive requests? What happens when the session expires? Who can take over for a second factor? Where does the run receipt go?
Make failure stop in one named place.
Chrome DevTools Recorder can add selectors, wait-for-element steps, and assertions for attributes, JavaScript properties, and visibility. Its documentation says a failed assertion reports an error after a timeout.[4]
That gives recorded automation a clean shape. Target the element by meaning. Wait for a real condition. Keep identity scoped. End with a witness that can fail loudly. The run should stop at the first broken contract instead of improvising through a portal with live authority.