Pimp My IDE / garage dispatchBack to dispatches
Model drift / longitudinal evals / 29 SEP 2026

A vibe is not a drift verdict.

When a hosted model feels worse, save the complaint. Then freeze the test, pin the harness, measure uncertainty, and wait for the second window.

The signal worth reading

Someone started the clock

Livenerf is testing a complaint that usually arrives as folklore. It runs a frozen panel against Claude Opus 5.5 once a day through a pinned Claude Code command-line harness. The project keeps raw logs, rejects samples touched by retries, refusals, or a different served model, and compares later windows with its launch-period baseline.

The important news is not a regression. There is no regression result yet. On September 29, the repository reported 6 of 30 days collected. Its pre-registered rule needs two consecutive 10-day windows in the same direction, a 99 percent interval that excludes zero, at least a 3-point change, an unchanged harness, and an error rate below 5 percent.

That restraint is the product. A complaint can choose the question. It cannot choose the verdict.

The rig measures the model as served through one named harness. It does not isolate a hidden raw model from routing, policy, or product behavior.
01 / BASELINE

Start before the rumor

Freeze prompts, items, scorers, effort, and schedule before later output can steer the test.

02 / HARNESS

Pin the delivery path

Record the client, version, settings, tools, model route, and code that shapes each sample.

03 / CONTROL

Prove the gauge moves

Use a positive control and an A/A check. A quiet gauge may be insensitive, not reassuring.

04 / DECISION

Publish no-call states

Set effect size, uncertainty, window, exclusions, and attribution rules before reading the result.

Interactive makeover / model drift witness rig

Make the claim cross every brake

A satisfaction poll can tell you where to look. This teaching rig shows why two noisy windows do not become a drift finding until effect size, uncertainty, direction, harness identity, and error limits agree. It mirrors livenerf's published rule. It does not run a model evaluation.

BASELINEWINDOWSHARNESSDECISION

Window controls

Move the observed deltas and their standard errors. The two window needles share one scale.

Measured windows
Decision interlocks
NO CHANGE DETECTEDTeaching replay

The dip is inside the noise brake.

Neither window clears the 99 percent uncertainty rule. Keep collecting and publish the no-call state.

What the gauge can miss

A null result has a size

Livenerf estimates that one 10-day window can detect a change of about 7.5 accuracy points at 80 percent power under its 99 percent test. A smaller real change may pass through the rig and still produce "no change detected." That sentence is different from "the model did not change."

The repository also found that lowering effort changed output tokens more clearly than accuracy during validation. Yet token count remains a secondary signal in the registered decision. This keeps a measurement discovered during setup from quietly replacing the declared primary metric.

The reusable move is simple. Log the serving path beside the score. Keep a control arm. Measure the gauge with a known change. Publish misses, exclusions, uncertainty, and the smallest change the run was built to see. Then let the calendar finish the test.

Source log / read 29 SEP 2026

What was read and what it supports

Artifact check: the repository exposes source, a committed pre-registration, frozen panel data, tests, append-only result conventions, and a live progress log. This pass read the public files but did not install the project, spend a subscription quota, or reproduce its statistics. The article therefore reports the project's methods and stated measurements. It does not endorse a drift finding, because the project has not made one.