Pimp My IDE / garage dispatchBack to dispatches
Decision systems / inference policy / 29 SEP 2026

Reasoning needs a clutch.

Jeeves publishes three useful operating points for one decision model. The fast path, bounded path, and full-thinking path differ by seconds and measured accuracy. Treat that choice as release policy.

The useful part

One model, three operating policies

PostHog released Jeeves as a 9B decision model that can answer yes or no, multiple-choice, and rating questions. Its request options can disable reasoning, cap the reasoning chain, or skip reasoning above a no-think confidence threshold.

The repository reports a clear trade. On 325 development questions served on one H100, no thinking reached 0.775 accuracy at about 0.3 seconds. A bounded setting reached 0.806 accuracy with 2.0-second median latency and 5.6 seconds at the 90th percentile. Full thinking reached 0.825 accuracy with a 3.3-second median and 17.1 seconds at the 90th percentile.

That table does not pick your production mode. The sample is the publisher's development set. Your class balance, failure cost, hardware, prompts, and traffic pattern can change the answer.

The practical move is to make reasoning mode part of the test matrix. Compare modes on the same holdout. Keep an abstain or human-review route. Measure tail latency beside decision quality.
Interactive makeover / reasoning clutch test bay

Select the lane. Build the test.

A normal model toggle hides the cost and evidence behind one switch. This clutch shows the repository's measurements and writes a workload-specific test template. It does not claim that the selected mode passed your test.

Inference selector

Choose one published operating point. The carriage, metrics, status, and packet share the native radio value.

Reasoning mode
Bounded lane selected
Evaluation packet sections

Published reference

The bounded lane spends less time at the tail in the publisher's dev run. Test whether its quality loss matters on your workload.

Accuracy0.806dev questions
Median2.0 slatency
p905.6 slatency
0 of 4 sections selectedNo local evidence is attached.
Read the labels

What the numbers do not settle

The public Jeeves results are broader than one win. The model card says Jeeves trails Jev on MMLU and MMLU-Pro. The GitHub table also shows Jev ahead on the combined MMLU-Pro and buried-state transfer row. A stronger total on another split does not erase that result.

The no-thinking and full-thinking scores also differ between the test split and the 325-question serving table. Keep those evaluations separate. Do not splice the best score from one table to the best latency from another.

The released weights require the Jeeves code because the pointer head and prompt format are part of the model. The code is MIT licensed. The derived weights use Apache 2.0. The model card states that distinction directly.

Benchmark before threshold

A confidence threshold can save work. Calibrate it on the traffic that will cross the route.

Keep a cheap baseline

Compare against a simple classifier and the no-thinking path. More reasoning has to earn its latency and operating cost.

Replay costly misses

Save false accepts, false rejects, and abstentions. Run the same cases after every prompt, model, or threshold change.