Cloudflare released Clef plus Clef-flash on 1 October, two open weight decision models under Apache 2.0. They take a state and up to 64 typed questions, then return probabilities rather than text. Clef-flash answers in 38.8ms median against Jev’s 524.1ms. Both read images. Both speak Jev’s API directly, so swapping is a config change.
Cloudflare trained them with a Brier loss, which exists for one purpose: making the returned probabilities trustworthy. Then it published no reliability data. Neither has anyone else in this category. One early tester got 0.04 from one model plus 0.54 from the other on the same question, which is what unchecked confidence looks like in practice. The 38.8ms figure is also text only. Image requests took 13 to 30 seconds on launch day.
Best for routing, classification plus scoring work where you want numbers your code can read. Not ideal for knowledge questions or anything where a wrong confidence score costs money.
A regular if statement works when your code can compute the condition.
It falls apart the moment the condition needs understanding. Is this support ticket angry. Does this message contain reproduction steps. Which team should handle this. You cannot write that in Python, so for two years the answer has been to ask a language model, get a paragraph back, then parse the paragraph and hope.
Decision models exist to delete the parsing step. You hand over a state plus a question with a fixed set of allowed answers. You get back a number.
Cloudflare just shipped two of them with the weights attached.
What Clef Actually Is
Here is the verified spec, taken from Cloudflare’s changelog, the model cards plus independent coverage on launch day.
| Item | Detail |
|---|---|
| Released | 1 October 2026 |
| Models | Clef at 27B, Clef-flash at 9B |
| Base models | Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash |
| Licence | Apache 2.0, open weights |
| Weights | Cloudflare/clef plus Cloudflare/clef-flash |
| Hosted | Workers AI |
| Clef pricing | $0.24 per million input tokens |
| Clef-flash pricing | $0.09 per million input tokens |
| Output pricing | None, since no text is generated |
| Context window | 65,536 tokens |
| Questions per request | Up to 64 |
| Images per request | Up to 4, maximum 4 MiB each |
| Request body cap | 13 MiB |
| Clef-flash latency | 38.8ms median |
| Clef latency | 209.3ms median |
| Jev latency for comparison | 524.1ms median |
| API | Speaks Jev’s SystemOne format directly, no shim |
| Calibration training | Rank-256 LoRA plus a Brier loss term, then an RL stage Cloudflare calls RLCD |
| Published calibration metrics | None |
Every figure there came from a primary page or launch day coverage rather than a summary of a summary.
Why a Decision Model Exists at All
The pitch is narrower than it first sounds, which is the reason it works.
A language model is built to produce text. When you want a decision from one, you write a prompt asking for JSON, then you validate the JSON, then you handle the case where it returned prose instead, then you handle the case where it invented a category you never offered. Everybody building agents has written that code. Everybody has watched it fail at three in the morning.
A decision model inverts the contract. You declare the answers up front. The model cannot return anything else, because returning something else is not a shape the output supports. Your code reads a float.
That constraint buys speed too. We covered the same trade in our comparison of automation platforms, where the tools that constrained what a step could do ran faster than the ones that allowed anything. Narrowing the output space narrows the work.
Clef pushes that further than most. It does a prefill only pass over your state, then a joint schema head scores every option of every question in parallel. No token by token generation happens at all. That is why a 9B model answers in under 40 milliseconds while a conversational model of similar size takes several hundred.
What that looks like in a real pipeline
Picture a support inbox. A ticket arrives, then your code needs four things before it can act. Is this urgent. Which team owns it. Does it contain enough detail to work on. How annoyed is the person writing it.
With a language model that is one prompt asking for structured output, a validation layer, a retry path plus a fallback for when the model invents a fifth team that does not exist. Latency lands somewhere between one and four seconds depending on how much it decided to explain itself.
With Clef it is one request carrying four typed questions. Urgency comes back as a probability. Team comes back as a distribution across the teams you declared. Detail sufficiency comes back as a boolean probability. Annoyance comes back as a position on a scale you defined. Total time under 40 milliseconds on the flash model, with no parsing code anywhere.
The request limit is 64 questions, so a realistic pipeline asks far more than four. You can score an entire triage decision tree in one call.
The Three Questions It Can Answer
Three question types cover the whole surface, which keeps the API small.
noul, for yes or no
You ask a boolean. You get a single probability between 0 and 1.
Does this support ticket describe a blocked customer. The answer comes back as 0.73. Your code decides what threshold means blocked.
choice, for picking one
You supply between 2 and 255 named options. You get a probability for every one of them.
Which team handles this ticket. Billing at 89%, technical at 11%. Note that you get the full distribution rather than just the winner, so your code can detect the case where two options are close plus escalate to a human.
score, for ordered scales
You supply between 2 and 10 ordered levels. You get a probability weighted position on that scale.
How severe is this issue, where the levels run cosmetic, workaround exists, then blocking. An answer of 1.3 sits between the second and third level, leaning toward workaround exists. That interpolation is more useful than a single bucket, because it tells you how close the call was.
All three take images. Up to four per request at 4 MiB each, which is what the vision encoder is for. Hold that thought.
Unless You Send an Image
Now the first caveat, which you will not find in the launch post.
That 38.8ms figure is text only. One tester working through the models on launch day reported image requests taking 13 to 30 seconds. Not milliseconds. Seconds.
Consider what that means for a product. The headline feature of these models is latency low enough to sit inside a request path, deciding how to route something while a user waits. The other headline feature is that they read images. On launch day you could use either one. Using both put you three orders of magnitude outside the number that made the models interesting.
That gap will probably close. Image preprocessing on a brand new endpoint is exactly the thing that gets optimised in the first month. The point for now is that two advertised capabilities currently cancel each other out, plus nobody announcing the models mentioned it.
If your use case is visual, benchmark it yourself this week before you design around it. Text requests are quick. Image requests on launch day were not.
The Confidence Score Nobody Has Checked
Here is the question VU has been chasing since the first of these models shipped.
Cloudflare did not simply train a classifier and read off the softmax. It added a rank-256 LoRA plus a Brier loss term, then ran a reinforcement learning stage it calls RLCD for ordinal choices. That Brier loss is the signal to read. The function exists for one job.
What a Brier loss is for
Calibration means your confidence numbers mean what they say. If a model outputs 0.80 across a thousand cases, roughly 800 of those should turn out true. A model can be accurate while being badly calibrated, saying 0.99 about things it gets right 70% of the time. For a chatbot that hardly matters. For code that branches on a threshold, it is the whole ballgame.
A raw softmax score is not a probability in that sense. It is a number between 0 and 1 that correlates loosely with correctness. Brier loss penalises the gap between stated confidence plus observed outcome, so training against it is a direct statement that the team wanted those numbers to be real.
Measuring whether it worked takes a reliability diagram. You bucket predictions by stated confidence, then plot observed accuracy against stated confidence, then look at how far the line sits from the diagonal. It is standard. It fits on one chart.
Cloudflare published accuracy benchmarks, latency medians plus pricing. No reliability diagram. No expected calibration error. Nothing.
Why that is stranger than it sounds
Every decision model on the market sells calibrated confidence as the core feature. Not one of them has published the measurement.
That was easier to overlook when the leading option was closed, since you simply could not check. Clef changes the situation, because Apache 2.0 weights mean anyone with a GPU plus a labelled dataset can compute the curve themselves. For the first time the question is answerable from outside.
There is already a hint of what the answer might look like. One tester asked both models whether a ticket contained reproduction steps. Clef said 0.04. Clef-flash said 0.54. Same family, same input, same question, one answer near certain and one a coin flip.
To be fair to Cloudflare, that tester concluded the question was vague rather than the models being wrong, which is a reasonable read. Vague questions produce vague answers, plus the guidance to re-read your own question before blaming the model is good advice. Yet it illustrates the problem exactly. Without a reliability curve you cannot tell a vague question from an overconfident model, because you have no baseline for what these numbers are worth.
Where It Wins and Where It Loses
The benchmark table is unusually honest, so read the losses alongside the wins.
Clef leads 7 of 10 decision benchmarks. On BANKING77, an intent classification set, it scores 94.20 macro-F1 against Jev’s 79.74. On CLINC150 with out of scope detection it scores 97.43 against 89.27. Those are not small margins. For routing plus classification work this is a different class of tool.
Then look at knowledge. GPQA Diamond: Clef 48.0, Jev 78.3. MMLU-Pro: Clef 65.9, Jev 82.7. Thirty points down on one, seventeen on the other.
That split is the useful content in the whole release. Clef is a machine for making bounded judgements about material you hand it. It is not a machine for knowing things. Ask it which of five teams should own a ticket plus it will beat everything in its class. Ask it a graduate level physics question plus it will do worse than a model a third its size that was built for recall.
Cloudflare also reports beating Jev in 3 of 4 workflow evaluations, which sits between those two poles and is the figure most people will actually feel.
None of this has been independently reproduced yet, so treat the table as vendor reported. That is the same standard we applied to Hermes Agent’s open source claims plus to every launch where a lab ran its own evals.
It Is Not Cheaper
Open weights plus lower latency reads like a cost win. Check the per token maths before you assume it.
Clef runs $0.24 per million input tokens on Workers AI. Clef-flash runs $0.09. There are no output charges, since nothing is generated, which does simplify your billing.
Per token, Clef costs roughly 5.7 times more than Jev. Speed went up plus the weights opened, while the unit price went up too.
Whether that lands as cheap depends entirely on your volume plus your tolerance for errors. Accuracy critical decisions at modest volume favour Clef easily, because one avoided mistake outweighs the token difference. Bulk classification at scale is where the multiplier starts to hurt, plus that is exactly the workload people reach for decision models to handle.
The escape hatch is the licence. Apache 2.0 weights mean you can self host, at which point you are paying for GPUs rather than tokens plus the comparison changes completely. A 9B model is small enough to run on hardware a small team can afford, which is the same calculation we walked through when pricing out Claude Pro against the API.
Where the crossover sits
Run the arithmetic for your own volume before deciding. Token pricing wins while your request count stays low, because a GPU sitting idle still costs money every hour. Self hosting wins once you are classifying continuously, since the marginal request becomes free. The crossover point depends on your hardware plus your traffic shape, so there is no universal answer here. What matters is that Apache 2.0 gives you the option at all, which a closed API does not.
Worth saying that Clef being built on Qwen3.8-27B plus Clef-flash on Qwen3.5-9B puts an interesting dependency underneath all of this. Qwen has been through one licence change already, which did not affect weights people had already downloaded. Apache 2.0 on these checkpoints is yours regardless of what happens upstream later.
That dependency cuts both ways though. Building on an existing open model is why Cloudflare could ship two sizes at once plus price them this low, since the expensive pretraining was already done. It also means the ceiling on these models is partly set by decisions made elsewhere. If the base family stops shipping open weights, the next version of Clef has a harder problem than this one did. Nothing suggests that is happening. It is simply the kind of structural dependency people notice only after it breaks.
What You Should Actually Do
If you are routing tickets, tagging content, scoring submissions or doing anything where a language model currently returns JSON you then validate, try Clef-flash this week. The API accepts the same request shape as Jev, so if you already have that integration it is a config change rather than a rewrite. Start with choice questions, since that is where the benchmark margins are largest.
If your workload is visual, test the image path before you plan around it. Launch day numbers were 13 to 30 seconds. Confirm whether that has improved rather than trusting the 38.8ms headline, which was measured on text.
If you are shipping anything where a wrong confidence score costs real money, build your own calibration check. Take a few hundred labelled examples from your own domain, run them through, bucket by stated confidence, then compare against what actually happened. That is an afternoon of work plus it tells you something nobody has published. Open weights mean nothing stops you.
If you are picking between Clef plus Jev, the deciding question is not speed. It is whether your decisions need knowledge or only judgement. Judgement about material you supply goes to Clef. Questions needing recall go elsewhere.
And if you run a lot of volume, price the self hosted path. A 9B Apache 2.0 model on your own GPU removes the token multiplier entirely.
Before you commit either way
One more check first. Run the same twenty questions through Clef plus Clef-flash side by side, then compare the probabilities. The two disagreed sharply for at least one tester, so knowing whether the smaller model tracks the larger one on your specific question set is worth an hour of your time. If they diverge on your data, the cheap model is not a cheap version of the expensive one. It is a different model wearing the same API.
The Part Worth Keeping
Cloudflare did the harder thing here. It open sourced the weights, named its base models, published benchmarks including the ones where it loses badly, documented its training objectives down to the loss function, then priced everything in public. Compared with how most model launches go, this is close to exemplary.
Which makes the missing chart more conspicuous rather than less.
A team that adds a Brier loss has told you it cares whether the confidence numbers are honest. The measurement that proves it worked fits on a single plot, uses a standard method plus would take an afternoon. Every company in this category has skipped it. Every one of them sells calibrated confidence as the reason to buy.
For the first time that gap is somebody else’s to close. The weights are Apache 2.0 and sitting on Hugging Face right now. Whoever posts the first reliability diagram for a decision model gets to define how the entire category is evaluated from here.
That matters more than it sounds, because this category is filling up fast. A competing decision model shipped the same week claiming a lead on a third party scoring index. Open weight clones of the closed option have been appearing for weeks. Benchmarks keep multiplying while the one measurement that would separate a good decision model from a confident one stays unpublished by everybody.
Accuracy tables tell you how often a model is right. A reliability diagram tells you whether to believe it when it says it is sure. For software that branches on a number, the second question is the one your users feel.
Charts and Blocks
Latency and price against Jev
Where each model wins
All figures vendor reported by Cloudflare on 1 October 2026. None independently reproduced at time of writing. Macro-F1 for the classification rows.
