Pimp My IDE / garage dispatch
Back to garage
September 28, 2026 | local decision models / Jeff 0.2.0

The bigger tiny model lost the lap.

Jeff's 2B model wins the aggregate benchmark. Its 0.8B sibling wins all three published game loops. Model size is not a routing policy.

Pick the smallest model that survives your actual workload, including the repeated mistakes that a one-step score hides.

A useful model can be narrow on purpose.

Jeff is a new set of Qwen3.5 and Gemma 4 fine-tunes for zero-shot classification. A request describes a state and supplies options. The model returns probabilities in one forward pass. It does not generate an answer token by token. The repository exposes a TypeSafe-compatible endpoint for choice, yes or no, and score questions.[1]

The 0.8B model is a real downloadable artifact. Its Hugging Face record lists 852,985,920 parameters, Apache 2.0 weights, a model revision, configuration files, and about 1.84 GB of stored files. We verified the repository and model records. We did not download the weights or reproduce the performance claims.[2]

Use a classifier as a classifier. Do not promote it to planner because it answers quickly.

The aggregate score picked 2B.

The project reports 83.1 percent overall accuracy for Jeff-Qwen3.5-2B and 79.1 percent for the 0.8B model across five public benchmarks. The 2B model also leads on the separate JevBench hard tier, 53.3 percent to 47.6 percent. On that table, 2B is the obvious choice.[1]

The same report warns that its comparison with Jev uses different samples of the same benchmarks. It also says the benchmark panel helped shape the format of at least half of each training family, even though panel items were excluded. Those disclosures do not erase the result. They define what the result can support.

The loops picked 0.8B.

The game harness describes legal moves and their consequences in words. Across 20 seeded episodes, the 0.8B model reports 6.55 Doom kills against minus 0.9 for 2B, 10.3 Frogger crossings against 6.0, and 57 Pac-Man pellets against 41.2. The larger model loses every published game comparison between the two Jeff variants.[1]

The authors think the 2B base model is more risk averse and say training may have made that behavior worse. That is a project interpretation, not a general law about parameter count. The concrete result is narrower. A model that scores higher over independent questions can still perform worse when occasional choices alter the next state.

Latency also belongs to the workload.

On the same RTX PRO 6000, the project reports 22 milliseconds per decision for 0.8B and 24 milliseconds for 2B. On an Apple M4 Max with MLX, it reports 28 milliseconds and 60 milliseconds. The CPU results are 463 and 708 milliseconds. The ranking stays the same, but the penalty changes with the machine.[2]

Do not blend these numbers with Jev's network-inclusive API timing. Hardware, runtime, request shape, and network path all belong on the test card. A fast median also says nothing about tail latency or memory pressure under the traffic you expect.

Make the selection test harder than the demo.

  1. Freeze the exact model revision, runtime, prompt format, option wording, and machine.
  2. Build a holdout from the decisions the application will make. Weight expensive errors separately.
  3. Run sequence tests when one answer changes the next state. Record recovery after a wrong turn.
  4. Compare a rule, a smaller model, and the preferred model under the same harness.
  5. Keep a fallback for low confidence, malformed input, and unfamiliar state.
Interactive makeover / model selection differential

Put the workload on the axle

Traditional purpose replaced: one sorted benchmark table. Better version: move between published workloads, keep both models on the same task-specific scale, then select the evidence a real release still needs.

Select the test road

The values below come from the Jeff project. They are not fresh measurements from this page.

Published workload
Required selection evidence
Selection draftAccuracy / higher is better
Published comparison

Workload differential

The published panel favors 2B.

One of four release requirements is selected. This comparison does not choose a production model.

What this component proves. It keeps six project-published comparisons attached to their workload and unit. It does not run either model, reproduce a benchmark, estimate calibration, or approve automated action.

Sources and limits

Open the source log
  1. Jeff repository, commit 985797fa2040d1c7a14822a96123c055800d5d09, read September 28, 2026. The README supplies the API shape, training description, benchmark table, game harness, game results, latency table, and project caveats. We cloned the 126-file repository and inspected its version and command entry points. We did not run the model.
  2. Jeff-Qwen3.5-0.8B model card and Hugging Face model API record, revision 987825634b36300f6eb2131e49b6d8a1b1205228, read September 28, 2026. These records verify that the weights, configuration, license, parameter count, and model files exist. Performance numbers remain project-published claims.
  3. AutoJev-27B repository, read September 28, 2026. Jeff credits this open recipe as its starting point. AutoJev documents one-pass probabilities, calibration, training, and its own hardware requirements. It does not verify Jeff's smaller models.
  4. Hacker News discussion for Jeff, item 49883844, September 28, 2026. The thread is the discovery route and contains practitioner questions about accuracy and use cases. Claims above come from the repository and model records.

Evidence boundary. The artifact exists, but this dispatch did not download the 1.7 GB weights or reproduce inference. Benchmark and latency values are the project's measurements. Game results use its harness, prompts, seeds, and episode count. The operating rule is our synthesis.