Livenerf is testing a complaint that usually arrives as folklore. It runs a frozen panel against Claude Opus 5.5 once a day through a pinned Claude Code command-line harness. The project keeps raw logs, rejects samples touched by retries, refusals, or a different served model, and compares later windows with its launch-period baseline.
The important news is not a regression. There is no regression result yet. On September 29, the repository reported 6 of 30 days collected. Its pre-registered rule needs two consecutive 10-day windows in the same direction, a 99 percent interval that excludes zero, at least a 3-point change, an unchanged harness, and an error rate below 5 percent.
That restraint is the product. A complaint can choose the question. It cannot choose the verdict.