A decision model is a narrower tool.
Ollaya runs open decision models on local hardware. A request contains text or JSON plus typed questions such as a choice, yes or no, or a score. The model returns answers and probabilities in one forward pass instead of generating text token by token.[1]
That shape fits routing jobs. Examples include classifying a support request, tagging a document, or choosing a queue. It does not make every model answer reliable. The project reports about 8 to 10 milliseconds for one five-question request on an RTX 4090. Its comparison with a hosted service includes network time, so the page labels the setups as different.[1]
Speed makes a decision cheap to ask. It does not make the decision cheap to get wrong.
A probability needs a local meaning.
Scikit-learn gives calibration a plain definition. Among predictions near 0.8, a well calibrated binary classifier should be correct about 80 percent of the time. The same documentation warns that one score can mix calibration, discrimination, and uncertainty. A lower loss does not prove better calibration by itself.[3]
That means a displayed 0.92 is not a universal safety grade. Test the model on your labels, your language, your traffic, and the mistakes your system can afford. Recheck it after the input mix or model changes.
Keep the fallback beside the threshold.
An independent decision-model benchmark compared one hosted decision model, constrained language models, and deterministic baselines under a frozen protocol. Results changed by task. The report publishes accuracy, calibration, latency, cost, failure modes, and raw logs. It also found a hard option-count boundary in one tested system.[4]
That benchmark did not test Ollaya. It supports a more general operating rule. Compare against a cheap baseline, test boundaries, and define what happens when the model refuses or lands below the threshold.
Wire four fuses before automation.
- Holdout: Freeze examples that match the task and its expensive mistakes.
- Calibration: Compare predicted probability with observed correctness by probability band.
- Abstain: Send low probability, malformed, and out-of-scope inputs to a named fallback.
- Replay: Save model, policy, input class, answer, probability, route, and later outcome.
The fusebox below writes a policy template. It does not run a model, estimate calibration, or authorize an automated action.