Pimp My IDE / Garage Dispatch
Back to garage
September 25, 2026 | local models / routing / probability

Confidence is not permission.

A fast local model can sort a ticket in milliseconds. The hard part is deciding when that answer may move work without a person.

The take: Put typed questions, a measured threshold, an abstain route, and replay evidence between a model probability and an automated action.
Wire a decision policy

A decision model is a narrower tool.

Ollaya runs open decision models on local hardware. A request contains text or JSON plus typed questions such as a choice, yes or no, or a score. The model returns answers and probabilities in one forward pass instead of generating text token by token.[1]

That shape fits routing jobs. Examples include classifying a support request, tagging a document, or choosing a queue. It does not make every model answer reliable. The project reports about 8 to 10 milliseconds for one five-question request on an RTX 4090. Its comparison with a hosted service includes network time, so the page labels the setups as different.[1]

Speed makes a decision cheap to ask. It does not make the decision cheap to get wrong.

A probability needs a local meaning.

Scikit-learn gives calibration a plain definition. Among predictions near 0.8, a well calibrated binary classifier should be correct about 80 percent of the time. The same documentation warns that one score can mix calibration, discrimination, and uncertainty. A lower loss does not prove better calibration by itself.[3]

That means a displayed 0.92 is not a universal safety grade. Test the model on your labels, your language, your traffic, and the mistakes your system can afford. Recheck it after the input mix or model changes.

Keep the fallback beside the threshold.

An independent decision-model benchmark compared one hosted decision model, constrained language models, and deterministic baselines under a frozen protocol. Results changed by task. The report publishes accuracy, calibration, latency, cost, failure modes, and raw logs. It also found a hard option-count boundary in one tested system.[4]

That benchmark did not test Ollaya. It supports a more general operating rule. Compare against a cheap baseline, test boundaries, and define what happens when the model refuses or lands below the threshold.

Wire four fuses before automation.

  1. Holdout: Freeze examples that match the task and its expensive mistakes.
  2. Calibration: Compare predicted probability with observed correctness by probability band.
  3. Abstain: Send low probability, malformed, and out-of-scope inputs to a named fallback.
  4. Replay: Save model, policy, input class, answer, probability, route, and later outcome.

The fusebox below writes a policy template. It does not run a model, estimate calibration, or authorize an automated action.

Interactive makeover / routing policy

Decision fusebox.

Traditional purpose replaced: a confidence cutoff hidden in application code. Better version: expose task impact, observed probability, threshold, fallback requirements, and a copyable policy on one control panel.

Set the route

Native controls own the state. Values are examples for writing a policy. They are not model measurements.

Decision lane
82%
0% uncertain100% model output
90%
0% open100% strict
Required policy fuses
HUMAN ROUTE0 / 4 FUSES SELECTED

Send this item to a person.

The example probability is below the policy threshold. No policy fuse is selected.

Decision laneTag
Route ruleHuman route below 90 percent
Policy claimDraft only. No task evidence is attached.
This teaching control compares two numbers and writes a template. It does not test a model, verify a reliability curve, detect drift, or make a production decision. The Act lane always keeps a person in the route.

Copy the policy

Fill after copyingAdd the exact model, revision, dataset, metrics, threshold rationale, fallback owner, and replay location.

Sources read, not vibes

Open the source log
  1. Ollaya project site: typed choice, yes or no, and score requests; local execution; project-published latency; TypeSafe-compatible endpoints; platform support; open model catalog; and the comparison caveat.
  2. Ollaya repository and v0.5.0 release: Rust implementation, Apache 2.0 license, model registry, command line and desktop artifacts, checksums, and current release assets. We downloaded the Linux amd64 release, matched its published SHA-256, and ran its help command.
  3. Scikit-learn probability calibration guide: operational definition of calibrated probability, reliability diagrams, scoring caveats, and cross-validation guidance.
  4. Decision Model Benchmark: independent frozen protocol, deterministic baselines, task-level accuracy, calibration, latency, cost, failure boundaries, and published raw logs. The benchmark tests Jev and constrained language models, not Ollaya.
  5. Hacker News discussion for Ollaya: exact discovery route and practitioner questions about hard-query performance. Comments are context, not independent model evidence.

Source boundary: Ollaya supports the product and compatibility claims. Its latency numbers come from the project and use named hardware. Scikit-learn supports the calibration definition. The independent benchmark supports the test method, but it does not measure Ollaya. The four-fuse policy is our synthesis.