Three weeks ago a startup called TypeSafe AI shipped Jev and called it a decision model. This week Cloudflare released two open-weight ones, and OpenAI previewed a Decisions API on top of GPT-6 Luna. The pitch is identical in all three: stop asking a frontier model to write a paragraph when you need it to pick one of five options, and get back a choice with a probability attached. The speed claims are checkable and are being checked. The probability claims are the point of the whole category, and almost nobody has published a number for them.
My read: decision models are a real idea and calibration is their entire product. A router that is fast and wrong with a confident 0.97 is worse than a slow LLM, because the confidence is what tells your code when to stop and ask a human. Right now that confidence is a marketing claim in two of the three launches and a measured one only in a test run by a company that sells calibrated confidence.
Latency gets audited. Calibration doesn’t.
Latency is the claim everyone can verify, so it gets verified. OpenAI’s DevDay slides say the Decisions API returns in 150 milliseconds against 1.6 seconds for a standard Luna call. Firecrawl’s comparison notes that figure comes with unstated conditions, while Jev’s numbers come from OpenRouter telemetry: 0.21 seconds at the median, 0.34 at p95, 0.75 at p99. Cloudflare’s Clef post reports 209.3 milliseconds median for Clef and 38.8 for the smaller Clef-flash, against 524.1 for Jev. Any of those can be reproduced with a stopwatch.
Calibration is the opposite. Cloudflare’s post says Clef was trained with label-smoothed cross-entropy and a Brier loss “to refine probability calibration.” It then reports accuracy and latency across 43 benchmarks and prints no calibration metric at all, no expected calibration error, no Brier score, no reliability curve, for Clef or for Jev. TypeSafe’s launch post says Jev “always communicates confidence and uncertainty” and, in the summary I fetched, offers limited empirical calibration data. OpenAI has published less: per Firecrawl, the Decisions API has no public contract, no schema and no pricing, and its probabilities are not documented. It is in limited preview.
Tim Fernholz put the open question plainly in TechCrunch: how well calibrated each of these models’ outputs will be to real life. TypeSafe’s own CEO, Diogo Almeida, said “Fast and cheap is very easy… Intelligence is the hard part.” He has an interest in saying so, and he is also right.
The economics explain why the gap matters. Per Firecrawl, Jev costs $0.042 per million input tokens and nothing for output, since it cannot emit text, while Luna is priced at $0.10 and $0.50. TypeSafe says its models run 40x to 200x faster than frontier models on these tasks and Fernholz relays its figure of $2.94 against $372 for the same monitoring workload. That is an overwhelming reason to switch, and an overwhelming reason to switch before anyone has checked whether the probabilities mean anything. A team that saves two orders of magnitude and loses calibration has not saved money. It has moved the cost to the incident report.

The one test that exists, and who ran it
The only independent-looking calibration data I could find is from Ryan Porter at Anthus, published October 1. He ran 3,600 problems from the ProofWriter logic dataset through GPT-6 Luna with a strict JSON schema and reasoning off. When Luna said it was 99% sure or more, it was right 68% of the time. At five inference steps its accuracy was 46% and 45% on the two task variants, and its AUROC, the measure of whether a high stated probability actually separates right answers from wrong ones, was 0.51. That is a coin. Jev on the same task scored 81% and 89% accuracy with an AUROC of 0.84, and was right 98.9% of the time when it claimed 99% or more.
That is a damning result for Luna as a classifier and a flattering one for Jev, and it deserves three caveats before anyone repeats it. First, Porter tested Luna directly, and says outright that this is not the Decisions API. OpenAI’s product is a constrained version of the model, and nothing published tells us how much that changes the confidence signal. Second, the data is synthetic and templated logic, which Porter flags himself as a limit on generalization. Third, Anthus lists calibrated confidence as a core offering, “which decisions to trust and which to escalate.” I found no disclosed financial tie to TypeSafe, and I am not suggesting the numbers are wrong. I am saying the single data point behind the story comes from a vendor in the category, and the category has no neutral referee yet.
Porter’s practical advice is the part I would keep: check that the confidence separates right from wrong on your hardest cases, because if it does not, no recalibration rescues it. Overconfidence can be corrected with a mapping. A score that carries no information about correctness cannot.
The size of the gap matters more than the headline. At depth 5 Jev led Luna by 35 points on one variant and 45 on the other, and the AUROC difference, 0.84 against 0.51, is the difference between a gate you can set a threshold on and one you cannot. If Luna carried over to the real Decisions API unchanged, a “route to a human below 90%” rule would send some of the wrong answers to a human and let the same share of right ones through. My inference is that the constrained product will do better than raw Luna, since OpenAI built it for exactly this job. Nothing published says how much better, and the preview has no documented probabilities to test.
Where the newcomers lose
The category is not a clean sweep for the purpose-built models either, and the launch posts say so when you read the tables. Cloudflare’s own benchmark run has Clef behind Jev on When2Call, 72.37 against 80.97, on BRIGHT, 45.91 against 47.52, and on PhishNChips, 79.60 against 85.35. It also trails on agent trace observability, 68.5 against 71.6. Where it wins it wins big: 94.20 against 79.74 on BANKING77, 97.43 against 89.27 on CLINC150+OOS. Cloudflare ran these itself, on its own models, and the post states as much. A fine-tune that gains accuracy in one domain gives up general performance elsewhere, which Cloudflare concedes.
When2Call is the one that should bother agent builders. It tests whether a model knows when to call a tool and when to ask or decline, which is the decision an agent safety layer actually makes. The model Cloudflare released as the next step in the category is more than eight points behind on that benchmark. The same goes for the claim that bounded decisions fix runaway agents. The Fernholz piece reports OpenAI’s own monitoring runs a separate model watching for bad actions at what the article calls significant compute cost, and cites Jev doing the same monitoring for $2.94 against $372 with a frontier LLM. The cost half is plausible. Whether the cheaper monitor catches the same actions is the half nobody printed.

What an audit would take
The odd thing is how little it would cost. A position paper from Sanz-Guerrero and von der Wense, submitted September 22 to an EMNLP workshop, notes that standard calibration metrics need only two inputs per example: a confidence score and a correctness judgment. Every benchmark Cloudflare ran already has the second. Any decision-model vendor has the first. The metric could ship in the same table as accuracy. It is a position paper without large experiments, so it argues for the practice rather than proving it, but the arithmetic is not in dispute.
Black-box audits are getting cheaper too. Plaud and colleagues estimate true calibration error for binary classification using only the logit_bias parameter, one query per sample, for APIs that hide their output probabilities. That is binary tasks only, and it presumes the API exposes logit_bias, so it does not cover a multi-option decision endpoint as it stands. But the direction is clear: a customer can start measuring a vendor’s confidence without the vendor’s cooperation.
There is a reason this matters more for agents than for a spam filter. A decision model in an agent loop is rarely called once. It classifies the request, picks the tool, judges the risk of the call, and decides whether a human should approve it. Each of those outputs carries its own stated probability, and the errors are not independent. A system built on four calls that are each 95% right and 99% confident has no easy way to know which of the four to distrust, and the vendors have published no composed-pipeline calibration to tell us.
So the standard I would hold the category to is simple. Before a team routes anything consequential on a stated probability, the vendor or the team publishes expected calibration error and AUROC on that team’s hardest cases, not on a launch benchmark. Firecrawl’s author put the architecture rule in one line: a decision model can judge meaning, but authority has to live in your code. I would add its corollary. The probability can gate an action only after someone has shown it separates the right calls from the wrong ones.
Decision models will win a lot of the cheap, boring work agents do, and they should. But three launches in three weeks have produced a mountain of latency charts and one calibration test, run by an interested party on synthetic logic. The first vendor to print its reliability curve next to its accuracy table will have said something its competitors haven’t. Until one does, treat the number after the decimal point as a claim.

AI-generated editorial illustration · TemperatureZero · October 2, 2026
Keep reading the signal
Get the Daily Signal — a concise briefing on what actually matters in AI and the systems around it.
Subscribe FreeContinue the archive