Clef has a real artifact.
Cloudflare describes Clef as a 27 billion parameter multimodal decision model. Clef-flash has 9 billion parameters. Both accept state plus typed questions and return a probability for each allowed option instead of free-form text. Cloudflare released the weights under Apache 2.0 and published model files, code, tokenizer files, and usage examples on Hugging Face.[1][2]
That clears an important first gate. These are downloadable model artifacts, not terminal-shaped product mockups. It does not prove that either model fits your task, accelerator, latency target, or error budget.
The first place changes with the column.
Cloudflare's leaderboard reports a 61.21 overall Decision Index for Clef and 57.07 for Clef-flash. The board marks both sets of results as self-reported and says the upstream board has not reproduced them. It also reports missing benchmarks as zero and warns that Cloudflare latency was not measured on the board's hardware.[3]
The rows move when the question changes. Clef leads the overall index among the listed models. Clef-flash has the higher Tools score of the two Cloudflare releases. Some community models show lower median latency on the board's RTX Pro 6000. The Cloudflare rows do not report expected calibration error.
The table contains useful evidence. The rank compresses away the question you still need to answer.
Probability still needs a local check.
Scikit-learn defines a calibrated binary classifier with a concrete test. Among predictions near 0.8, about 80 percent should belong to the positive class. Its guide also warns that a single score can mix calibration, discrimination, and uncertainty. A lower score does not isolate calibration quality by itself.[4]
If a decision model will route work, collect reliability bands on your own labels. Keep malformed input and low-confidence cases on a named fallback. Recheck after the model, question schema, or input mix changes.
Use the board to write the test.
- Pick the exact task and costly error.
- Choose the task-level metric before reading the rank.
- Run candidates on matched hardware and the same request path.
- Measure calibration or reliability bands on held-out local examples.
- Replay the route with the abstain path and later outcome attached.
The differential below turns the published comparison into a review card. It does not select a model or reproduce the published scores.