Pimp My IDE / garage dispatch
Back to garage
October 1, 2026 | browser engines / generated code / conformance

A browser oracle can still teach the wrong lesson.

Bez is testing whether models can write parts of a web engine from specifications while three shipping browsers grade the geometry. The sharp idea is the witness loop. The dangerous shortcut is treating browser agreement as the specification.

Pin the rule. Prove the tests reject bad candidates. Compare independent implementations. Keep a held-out replay. Four different jobs, one release decision.

The interesting part is not the code generator.

Bez describes a pipeline that gives a model one specification section, typed Rust inputs and outputs, and failing examples. The model writes several candidates. Each candidate compiles inside the engine and runs against cached geometry from Chromium, Firefox, and WebKit. A candidate can land only when it keeps prior passing cases and beats the committed score.[1]

That is a better shape than asking a model to "write a browser." The model gets a narrow function. The harness owns the types, corpus, comparison, score floor, and provenance record. The accepted result becomes ordinary Rust. No model runs during the engine build or at runtime.[1]

Generation is cheap only after someone builds a test that knows how to say no.

Agreement is evidence, not authority.

The project's geometry check renders one page in three browsers at a fixed viewport. It records box rectangles and selected computed styles. When two browsers agree and one differs, the pair supplies the reference. When all three split, the measurement is skipped.[2]

This is a useful compatibility oracle. It answers whether a candidate matches the behavior developers already encounter. It cannot prove that the majority followed the written rule. Browsers share history, copied behavior, and compatible bugs. A specification citation and a browser vote belong on the same review sheet, but they do not replace each other.

The Web Platform Tests project makes the same cross-browser purpose explicit. Its suite helps implementations ship compatible behavior and gives developers confidence that web features work across browsers. WPT is an independent source of tests. It is not a promise that every platform rule has a complete test.[3]

Read the small numbers before the big thesis.

The repository's September 25 dashboard marks 0.6 percent of 17,259 compatibility keys as generated and 93 percent as unreached. Its README says nine CSS 2.1 layout rules sit in the generated directory. Eight were model-written. One stayed hand-written because no candidate beat it.[1]

The narrow corpus tells a second story. The project reports 227 of 227 recipe cases passing. Its geometry document also reports 61 of 164 grammar cases scored, with 36 discarded because browsers disagreed about scrollbars. The held-out normal-flow WPT slice is seven pages, not the full suite.[2]

Those limits make the work more credible, not less. The repository labels missing areas, rejected measurements, and open questions. A prototype earns attention when its dashboard can show the empty floor.

The hand-built harness carries the claim.

Bez says the expensive step is writing the function boundary, baseline, recipe corpus, and deliberately wrong candidates. Those wrong candidates matter. If the corpus accepts them, the test is not ready to grade generated code.[1]

The broader web platform makes this hard to scale. WHATWG's current standards index spans HTML, DOM, Fetch, Storage, Streams, Web IDL, and device-facing APIs. The W3C browser-specs catalog exists partly to connect specifications with machine-readable references and test paths.[4][5] A layout box can be compared numerically. Permissions, accessibility, timing, media, networking, and security UI need different witnesses.

The practical lesson reaches beyond browser engines. For generated parsers, compilers, database kernels, and protocol code, review the harness before admiring the output. Find the normative rule. Seed a known bad implementation. Keep a separate compatibility witness. Replay cases the generator never saw.

Interactive makeover / browser oracle roll cage

Make the witness chain touch

Traditional purpose replaced: one green conformance badge. Better version: four native controls expose the rule, rejection power, compatibility witness, and held-out replay. The rail advances only through the contiguous checked prefix.

Select the review sections

These controls build a review template. They do not run Bez, WPT, or a browser.

Ordered witness sections

A downstream selection does not repair an upstream gap. The active node lights, but the route stays broken at the first missing section.

Evidence route

Four-point roll cage

0 sections selected
RuleOpen
RejectionOpen
OracleOpen
ReplayOpen
Route blocked at rule

No review section is selected.

Start with the normative rule. Every field still needs real evidence from the candidate and its harness.

Four witnesses, four jobs

Do not merge their meanings.

01 / NORMATIVE

What should happen?

Pin the specification text and revision. A compatibility majority can disagree with the rule.

02 / ADVERSARIAL

Can the test refuse?

Run a known bad candidate. A corpus that cannot reject it is not an admission gate.

03 / COMPATIBILITY

What ships now?

Compare independent implementations. Record splits, discarded cases, tolerance, and platform conditions.

04 / HELD OUT

Does the lesson travel?

Replay an unseen fixture against the exact candidate. Keep generator feedback out of this lane.

Sources read

Source log and evidence boundary
  1. Bez repository and README, inspected October 1, 2026 at commit be130e57d74b1aa3d065e451d7c5e78c1cc454b2. The repository describes the generation loop, current dashboard, generated rule count, admission floors, provenance records, and open questions. A local clone resolved to that commit. This dispatch inspected the source tree and project documents. It did not build the Rust workspace because the inspection host did not have Cargo installed.
  2. Bez, "How layout output is checked against real browsers", inspected from the same commit. It defines the fixed viewport, three-browser vote, 1/60 pixel tolerance, recipe corpus, held-out pages, discarded scrollbar cases, and reported current results.
  3. Web Platform Tests documentation, read October 1, 2026. The project describes WPT as a cross-browser test suite and links its canonical source, public deployment, and archived browser results.
  4. WHATWG standards index and HTML Living Standard, read October 1, 2026. They show the active standards surface and connect HTML sections to ongoing web-platform tests.
  5. W3C browser-specs repository, read October 1, 2026. Its README describes the curated specification catalog and its use for machine-readable references, compatibility data, web-features, and test-path analysis.
  6. Hacker News item 49925036, fetched from the official API on October 1, 2026. It identified Bez during a scan that also covered Cloudflare Clef, K2, current Rust compiler work, and OpenDLSS. It is a discovery source, not evidence for the project claims.

Evidence boundary. Project percentages and test counts are the Bez maintainers' records at the inspected revision. Pimp My IDE cloned and read the repository, confirmed its current commit, and inspected the workspace manifest. It did not compile the engine, run its browser oracle, reproduce its measurements, or judge browser completeness. The roll cage builds a template only. Checked sections do not mean evidence was collected.