Pimp My IDE / Garage Dispatch
Back to garage
September 26, 2026 | coding agents / reference systems / regression proof

Give the rebuild a witness lap.

A Prince of Persia port improved when the coding agent could run the original game, inspect known source, and compare rendered output. The useful part is the loop, not the model name.

The take. When an agent rebuilds an existing system, connect it to a legal reference copy, a repeatable input sequence, a machine-readable comparison, and a replay test. Source alone may explain the design. The running reference shows what the design does.
Build a witness loop

The first ports could read code but missed the movement.

Priyan Rajeevan describes a series of AI-assisted attempts to rebuild the 1989 Prince of Persia in C#. The first version parsed rooms, tiles, gates, and guard positions. It still moved the prince one tile at a time. The original uses frame sequences with smaller movement steps. Later surface fixes did not repair that wrong design.[1]

This is one person's experiment, not a controlled model benchmark. The prompts, tools, and code changed between rounds. The lesson still travels: a port can parse the right data and render familiar shapes while missing the behavior that makes the system itself.

A rebuild needs a witness that can disagree with plausible code.

The agent gained a reference it could operate.

Rajeevan reports that later sessions used skills to control DOSBox, send keys, and capture screenshots from a personal copy of the DOS game. The agent compared its port against the running game. It then replaced the tile-grid movement design with a frame-sequence design and used data from the original game files at runtime.[1]

The published repository has a headless mode that writes one PNG per tick. It accepts scripted key sequences. Its README separates working behavior from unfinished work, including guards, sword fighting, sound, palace art, and exact landing positions. That boundary matters more than a broad claim that the game was ported.[2]

Prior art supplied the missing map.

The final drawing work did not come from the DOS binary alone. Rajeevan says the agent found SDLPoP and ported its room-drawing routine. SDLPoP describes itself as an open-source port based on a disassembly of the DOS version. Its documentation credits earlier technical work and Jordan Mechner's released Apple II source.[3]

Mechner's repository contains the original 6502 assembly source from 1985 through 1989. Its README says the archive was recovered from old floppy disks and published for study. It also states that the source release grants no rights in the Prince of Persia franchise.[4]

This makes the evidence chain specific. The Apple II source explains an original design. SDLPoP reconstructs DOS behavior. The operator's DOS copy supplies runtime data and screenshots. The new C# port can then compare its output against a reference. None of those sources can replace the others.

A tiny pixel count can hide a large test gap.

Rajeevan reports that one checked room fell from 8,429 differing pixels to two. The repository uses narrower wording: room drawing is pixel-exact on the rooms checked against DOSBox. It also says only one DOS copy has been tested. Those limits belong beside the result.[1][2]

A screenshot diff checks one input, one time, one capture path, and one part of the screen. It does not prove movement, collision, timing, sound, or every room. Each behavior needs its own input script and comparison rule.

Build the loop before ranking the driver.

  1. Reference. Name the exact version, legal access path, revision, assets, and known differences.
  2. Drive. Use one recorded input sequence against both systems. Pin timing and initial state.
  3. Compare. Choose the right oracle for the behavior. Use pixels for rendering, state traces for movement, and events for side effects.
  4. Replay. Save the inputs, outputs, thresholds, and failure image. Run them again after the next change.

The test track below writes a witness-loop plan. It does not run DOSBox, inspect game files, compare screenshots, or certify a port.

Interactive makeover / reference-test planner

Witness loop test track

Traditional purpose replaced: one "looks right" checkbox. Better version: choose the observed behavior, close four ordered test gates, and copy a plan that keeps real evidence fields blank until you fill them.

Set the comparison lane

The native radio group chooses one observation type. Each square gate adds a required section. A selected gate records a test requirement, not a passing result.

Observed behavior
Witness loop gates
TRACK OPEN0 / 4 GATES
Build lane / reference laneThe build stops at the first open gate. The reference stays pinned at the start until its identity is recorded.

Pin the reference first.

No test section identifies which running system the rebuild must match.

ReferenceOPEN
DriveOPEN
CompareOPEN
ReplayOPEN

Four selected gates means the test plan has four sections. It does not mean the rebuild matches the reference.

Print the witness plan

Replace each required marker with exact versions, commands, outputs, and results. Keep unsupported behaviors in the open list.

Sources read, not vibes

Source log
  1. Priyan Rajeevan, "Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia", September 25, 2026. First-person report of the staged port, DOSBox observation, use of SDLPoP, pixel comparisons, reversions, and remaining work.
  2. PrinceOfPersia_C-Sharp_Port_By_AI repository, read September 26, 2026. Published code, staged history, scripted headless capture, runtime data requirements, checked-room wording, tested-copy limit, credits, license, and unfinished behavior list.
  3. NagyD/SDLPoP, read September 26, 2026. Open-source DOS port based on disassembly, contributor record, technical lineage, added-feature boundary, and GPL terms.
  4. Jordan Mechner's Prince of Persia Apple II source archive, read September 26, 2026. Original 6502 source history, archive recovery, study purpose, and rights boundary.
  5. Hacker News discussion for the Prince of Persia experiment, GitHub Trending, and Claude Code skills documentation were inspected for current context. They do not support the port's measured results.

Source boundary. The model-by-model account and pixel counts come from the project's author. Pimp My IDE read the public repository and cited upstream projects. It did not run the DOS game, build the C# port, reproduce the screenshots, or compare model versions. The test track is a planning tool, not benchmark evidence.