PostHog released Jeeves as a 9B decision model that can answer yes or no, multiple-choice, and rating questions. Its request options can disable reasoning, cap the reasoning chain, or skip reasoning above a no-think confidence threshold.
The repository reports a clear trade. On 325 development questions served on one H100, no thinking reached 0.775 accuracy at about 0.3 seconds. A bounded setting reached 0.806 accuracy with 2.0-second median latency and 5.6 seconds at the 90th percentile. Full thinking reached 0.825 accuracy with a 3.3-second median and 17.1 seconds at the 90th percentile.
That table does not pick your production mode. The sample is the publisher's development set. Your class balance, failure cost, hardware, prompts, and traffic pattern can change the answer.