The narrow model is the interesting part.
Jeff fine-tunes Qwen3.5 and Gemma 4 models to choose among supplied options in one forward pass. It returns a probability for each choice instead of generating prose. The project reports 22 milliseconds for its 0.8B model on an RTX PRO 6000 and 28 milliseconds on an Apple M4 Max. Those are project measurements on named hardware.[1]
The released 0.8B model has a public Hugging Face card, an Apache 2.0 license, a pinned repository revision, and 37 listed files. The card labels it for zero-shot classification. That proves an inspectable artifact exists. We did not download the 1.7 GB weights or run the model.[2]
Use a small model as a choice engine, not a tiny general-purpose oracle.
The larger score lost the game lap.
Jeff reports 83.1 percent overall panel accuracy for its 2B Qwen model and 79.1 percent for the 0.8B model. The order reversed in the project's game harness. The 0.8B model beat the 2B model in its reported Doom, Frogger, and Pac-Man results. The authors say benchmark scores did not predict game play and note that an occasional wrong move compounds inside a real-time loop.[1]
This is not proof that smaller models are better. The game prompts, state descriptions, action consequences, and repeated dynamics differ from the five-benchmark panel. It is proof that a higher aggregate score does not settle a loop-shaped deployment.
Single-step accuracy shrinks across a chain.
The dyno below uses a simple teaching calculation. If each decision has the same independent chance of being correct, the chance that all decisions are correct is the single-step probability raised to the number of decisions. NIST's binomial reference states the assumptions plainly: two outcomes, a fixed success probability, and repeated trials.[4]
Real agent steps are rarely independent. One wrong action can alter every later input. Error rates also change by state, label, and prompt. The calculation is a warning light, not a forecast. Replay the actual sequence on a held-out workload.
Test the loop, not the badge.
- Keep a rule or random baseline in the same harness.
- Score complete sequences and costly failure states, not only isolated choices.
- Set a stop budget before the model enters the loop.
- Save the model revision, prompt shape, state, options, probability, action, and next state.
- Repeat the test after the model, wording, workload, or action set changes.
MicroLLM Lab points in the right direction by putting model choice, runtime readings, comparisons, and custom evaluations in one browser workbench. Its seven tiny language models are not Jeff and its chat tasks do not verify decision-model performance. The useful idea is the visible test surface.[5]