The interesting part is not the code generator.
Bez describes a pipeline that gives a model one specification section, typed Rust inputs and outputs, and failing examples. The model writes several candidates. Each candidate compiles inside the engine and runs against cached geometry from Chromium, Firefox, and WebKit. A candidate can land only when it keeps prior passing cases and beats the committed score.[1]
That is a better shape than asking a model to "write a browser." The model gets a narrow function. The harness owns the types, corpus, comparison, score floor, and provenance record. The accepted result becomes ordinary Rust. No model runs during the engine build or at runtime.[1]
Generation is cheap only after someone builds a test that knows how to say no.
Agreement is evidence, not authority.
The project's geometry check renders one page in three browsers at a fixed viewport. It records box rectangles and selected computed styles. When two browsers agree and one differs, the pair supplies the reference. When all three split, the measurement is skipped.[2]
This is a useful compatibility oracle. It answers whether a candidate matches the behavior developers already encounter. It cannot prove that the majority followed the written rule. Browsers share history, copied behavior, and compatible bugs. A specification citation and a browser vote belong on the same review sheet, but they do not replace each other.
The Web Platform Tests project makes the same cross-browser purpose explicit. Its suite helps implementations ship compatible behavior and gives developers confidence that web features work across browsers. WPT is an independent source of tests. It is not a promise that every platform rule has a complete test.[3]
Read the small numbers before the big thesis.
The repository's September 25 dashboard marks 0.6 percent of 17,259 compatibility keys as generated and 93 percent as unreached. Its README says nine CSS 2.1 layout rules sit in the generated directory. Eight were model-written. One stayed hand-written because no candidate beat it.[1]
The narrow corpus tells a second story. The project reports 227 of 227 recipe cases passing. Its geometry document also reports 61 of 164 grammar cases scored, with 36 discarded because browsers disagreed about scrollbars. The held-out normal-flow WPT slice is seven pages, not the full suite.[2]
Those limits make the work more credible, not less. The repository labels missing areas, rejected measurements, and open questions. A prototype earns attention when its dashboard can show the empty floor.
The hand-built harness carries the claim.
Bez says the expensive step is writing the function boundary, baseline, recipe corpus, and deliberately wrong candidates. Those wrong candidates matter. If the corpus accepts them, the test is not ready to grade generated code.[1]
The broader web platform makes this hard to scale. WHATWG's current standards index spans HTML, DOM, Fetch, Storage, Streams, Web IDL, and device-facing APIs. The W3C browser-specs catalog exists partly to connect specifications with machine-readable references and test paths.[4][5] A layout box can be compared numerically. Permissions, accessibility, timing, media, networking, and security UI need different witnesses.
The practical lesson reaches beyond browser engines. For generated parsers, compilers, database kernels, and protocol code, review the harness before admiring the output. Find the normative rule. Seed a known bad implementation. Keep a separate compatibility witness. Replay cases the generator never saw.