Pimp My IDE / garage dispatch
Back to garage
September 28, 2026 | small models / repeated choices / control loops

A benchmark win can lose the long race.

Jeff puts small, fast decision models on local hardware. Its own game tests show why a one-step score cannot certify a repeated loop.

Benchmarks grade isolated choices. Agent routes, game loops, and automation chains pay for every mistake that survives into the next step.

The narrow model is the interesting part.

Jeff fine-tunes Qwen3.5 and Gemma 4 models to choose among supplied options in one forward pass. It returns a probability for each choice instead of generating prose. The project reports 22 milliseconds for its 0.8B model on an RTX PRO 6000 and 28 milliseconds on an Apple M4 Max. Those are project measurements on named hardware.[1]

The released 0.8B model has a public Hugging Face card, an Apache 2.0 license, a pinned repository revision, and 37 listed files. The card labels it for zero-shot classification. That proves an inspectable artifact exists. We did not download the 1.7 GB weights or run the model.[2]

Use a small model as a choice engine, not a tiny general-purpose oracle.

The larger score lost the game lap.

Jeff reports 83.1 percent overall panel accuracy for its 2B Qwen model and 79.1 percent for the 0.8B model. The order reversed in the project's game harness. The 0.8B model beat the 2B model in its reported Doom, Frogger, and Pac-Man results. The authors say benchmark scores did not predict game play and note that an occasional wrong move compounds inside a real-time loop.[1]

This is not proof that smaller models are better. The game prompts, state descriptions, action consequences, and repeated dynamics differ from the five-benchmark panel. It is proof that a higher aggregate score does not settle a loop-shaped deployment.

Single-step accuracy shrinks across a chain.

The dyno below uses a simple teaching calculation. If each decision has the same independent chance of being correct, the chance that all decisions are correct is the single-step probability raised to the number of decisions. NIST's binomial reference states the assumptions plainly: two outcomes, a fixed success probability, and repeated trials.[4]

Real agent steps are rarely independent. One wrong action can alter every later input. Error rates also change by state, label, and prompt. The calculation is a warning light, not a forecast. Replay the actual sequence on a held-out workload.

Test the loop, not the badge.

  1. Keep a rule or random baseline in the same harness.
  2. Score complete sequences and costly failure states, not only isolated choices.
  3. Set a stop budget before the model enters the loop.
  4. Save the model revision, prompt shape, state, options, probability, action, and next state.
  5. Repeat the test after the model, wording, workload, or action set changes.

MicroLLM Lab points in the right direction by putting model choice, runtime readings, comparisons, and custom evaluations in one browser workbench. Its seven tiny language models are not Jeff and its chat tasks do not verify decision-model performance. The useful idea is the visible test surface.[5]

Interactive makeover / decision endurance dyno

Run the full choice chain

Traditional purpose replaced: one accuracy badge beside a model name. Better version: expose loop length, show the independence-based warning, close four test breakers, and copy a workload-specific endurance plan.

Set the test lap

The published presets load Jeff's reported panel figures. They are not measurements for your workload.

83.1%
50% teaching floor100% no observed misses
12
1 choice30 choices
Workload lane
Endurance test breakers
Test plan open1 of 4 breakers selected
Independence-based warning model

Cumulative clean-lap chance

10.8%Chance all 12 choices are correct if each has fixed 83.1% accuracy and decisions are independent
Expected wrong choices2.03
Chance of at least one miss89.2%

The loop needs three more test sections.

One domain holdout section is selected. The calculation is a teaching warning, not workload evidence.

What this component proves. It performs a binomial teaching calculation under fixed-probability and independence assumptions. It writes a test-plan structure. It does not measure a model, predict a real control loop, or approve automation.

Sources and limits

Open the source log
  1. Firelex, Jeff repository, read September 28, 2026. The README supplies the architecture, model sizes, project-measured latency, benchmark panel, game harness results, training notes, usage guidance, and caveats. Jeff's authors say their 2B model scored higher on the panel but played the three games worse than the 0.8B model.
  2. Jeff-Qwen3.5-0.8B model card and repository, read September 28, 2026. The Hugging Face API returned revision 987825634b36300f6eb2131e49b6d8a1b1205228, the zero-shot-classification task tag, the Transformers library tag, and 37 listed files. A request for config.json resolved against that revision. We did not download or execute the weights.
  3. Denis Yarats, AutoJev-27B repository, read September 28, 2026. The project documents a larger open decision model, one-forward-pass choice probabilities, separate temperature calibration, its published evaluation, and a roughly 49 GiB BF16 runtime requirement. Jeff builds on this open recipe.
  4. NIST/SEMATECH e-Handbook, Binomial Distribution, read September 28, 2026. It defines repeated binary trials with a fixed single-trial success probability. The dyno adds the explicit independence assumption needed for multiplying clean-step probabilities.
  5. MicroLLM Lab, read September 28, 2026. The live browser workbench lists seven quantized tiny language models, WebGPU and fallback runtime details, benchmark comparison, certificates, and custom evaluation. It is a useful interface reference, not evidence about Jeff.
  6. Hacker News discussion for Jeff, September 28, 2026. Used for discovery and practitioner reaction only.

Evidence boundary. Jeff's benchmark, speed, training, and game figures are project-published results. The compared systems and samples are not all matched. This article did not reproduce them. The dyno's cumulative values assume fixed independent trials, which real decision loops often violate.