Pimp My IDE / garage dispatch
Back to garage
October 2, 2026 | decision models / benchmarks / routing

A leaderboard is not a routing policy.

Cloudflare released two open-weight decision models and put Clef at the top of a published Decision Index. The same table also shows why one rank cannot choose your production route.

Read the task column, metric, hardware, artifact, and missing evidence before the top row earns a job.

Clef has a real artifact.

Cloudflare describes Clef as a 27 billion parameter multimodal decision model. Clef-flash has 9 billion parameters. Both accept state plus typed questions and return a probability for each allowed option instead of free-form text. Cloudflare released the weights under Apache 2.0 and published model files, code, tokenizer files, and usage examples on Hugging Face.[1][2]

That clears an important first gate. These are downloadable model artifacts, not terminal-shaped product mockups. It does not prove that either model fits your task, accelerator, latency target, or error budget.

The first place changes with the column.

Cloudflare's leaderboard reports a 61.21 overall Decision Index for Clef and 57.07 for Clef-flash. The board marks both sets of results as self-reported and says the upstream board has not reproduced them. It also reports missing benchmarks as zero and warns that Cloudflare latency was not measured on the board's hardware.[3]

The rows move when the question changes. Clef leads the overall index among the listed models. Clef-flash has the higher Tools score of the two Cloudflare releases. Some community models show lower median latency on the board's RTX Pro 6000. The Cloudflare rows do not report expected calibration error.

The table contains useful evidence. The rank compresses away the question you still need to answer.

Probability still needs a local check.

Scikit-learn defines a calibrated binary classifier with a concrete test. Among predictions near 0.8, about 80 percent should belong to the positive class. Its guide also warns that a single score can mix calibration, discrimination, and uncertainty. A lower score does not isolate calibration quality by itself.[4]

If a decision model will route work, collect reliability bands on your own labels. Keep malformed input and low-confidence cases on a named fallback. Recheck after the model, question schema, or input mix changes.

Use the board to write the test.

  1. Pick the exact task and costly error.
  2. Choose the task-level metric before reading the rank.
  3. Run candidates on matched hardware and the same request path.
  4. Measure calibration or reliability bands on held-out local examples.
  5. Replay the route with the abstain path and later outcome attached.

The differential below turns the published comparison into a review card. It does not select a model or reproduce the published scores.

Interactive makeover / decision index differential

Shift the question.

Traditional purpose replaced: one sorted leaderboard. Better version: the native selector changes the metric, keeps source provenance visible, and builds a copyable matched-test plan.

Select the lane

Values below come from Cloudflare's published leaderboard. Orange rows are Cloudflare self-reports. Blue rows come from the upstream board.

Comparison question
Matched-test sections to request
Published comparison

Overall index

0 of 4 test sections selected

Chance-corrected aggregate from Decision Index 0.2.1. Higher is better.

Cloudflare self-reportedUpstream board

Clef leads this published aggregate.

The source labels stay attached. A local winner still needs a matched test on your workload.

This is a teaching view of selected published rows. It is not a full leaderboard, independent reproduction, model run, calibration test, or deployment recommendation.

Matched-test card

The all-selected state means the test sections are requested. It does not mean a model passed.

Sources read

Source log and evidence boundary
  1. Cloudflare, "Introducing Clef", read October 2, 2026. This supplies the release claims, model roles, selected benchmark rows, Cloudflare's domain-classification example, Workers AI route, reinforcement-learning product announcement, and Apache 2.0 claim.
  2. Cloudflare Clef-flash model card and the Hugging Face model API, read October 2, 2026. These verify the 9 billion parameter artifact, Apache 2.0 metadata, file inventory, typed-question interface, custom model code, model revision, and single-H200 usage note. The Clef repository verifies the larger artifact.
  3. Decision Model Leaderboard, read October 2, 2026. This supplies Decision Index 0.2.1 values, category scores, source labels, hardware caveat, missing-benchmark rule, and the warning that Cloudflare's Clef results are self-reported and not reproduced by the upstream board.
  4. Scikit-learn probability calibration guide, read October 2, 2026. This supplies the operational calibration definition, reliability-diagram method, and warning that one score can mix calibration, resolution, and uncertainty.
  5. Hacker News discussion for the Clef release, item ID verified through the Hacker News API and read October 2, 2026. This is the discovery and practitioner-discussion route, not technical authority.

Evidence boundary. Pimp My IDE did not run either model or reproduce the benchmark. The interactive differential uses selected values from Cloudflare's leaderboard. Latency rows preserve its warning that Cloudflare and upstream measurements used different test cells. Missing calibration values remain missing instead of being estimated.