Do You Even Need a Decision Model?
A field test of 15 decision models on fresh data: when the new model class pays off, when a simple baseline is enough, and what is left of the marketing.
TL;DR. Decision models are AI models that do not write text. Given a question, they pick one of several predefined answers and return a probability. Since TypeSafe launched Jev in September 2026, many of them have appeared. I tested 15 on fresh data.
- Open models are catching up: Cloudflare’s Clef, Perplexity’s pplx-decider and APUS 9B are practically as accurate as Jev and freely available.
- Accuracy barely separates them: Six models cannot be told apart from Jev. Their prices differ by up to a factor of 13.
- Same interface, different behavior: A field called “confidence” is not automatically something you can rely on. Models also behave very differently on long texts and on many cases per request.
- How much can run unattended: Equally accurate models differ widely in how many cases they could take over without human review. Jev does worst here.
- With labels: If you already have labeled examples, a simple baseline of embeddings plus logistic regression is your best bet. It even beat fine-tuned decision models.
- Marketing: Many of the terms around Jev are new names for familiar things.
Transparency note: This text was written with AI, and AI agents ran the experiments. I planned, supervised and checked both. How that worked is described at the end. Code, data and every model response are in the linked repository.
On September 15, 2026, the startup TypeSafe presented Jev as the first “System One” model, one that decides instead of writing. Two weeks later, plenty of models shared the same interface, one of them from Cloudflare. This post tests 15 of them and answers a simple question: do you even need a model like this?
What a decision model is, and what part 1 showed
A decision model gets a text and a question with fixed answer options, for example: “Which department should handle this ticket?” It picks an answer and says how confident it is. The technical term is zero-shot classification: the model assigns a text to a category without having seen examples of it first. The idea is not new. It has been around since 2008.
In part 1 I tested Jev on fresh data: 400 scientific papers from September 2026, each to be assigned to one of eight computer science categories. The main results:
- Against small language models such as GPT-6 Luna, Jev was neither measurably better nor worse, but cheaper and faster.
- Against open models running on a laptop, Jev was clearly ahead. Only small models that fit into 16 GB of memory were tested, though.
- Jev’s confidence values were right on average, but for half of the answers Jev reported a confidence of exactly 100%. Among those answers, there is no way left to sort out which ones a human should check.
- With labels, meaning labeled examples, a simple baseline caught up: embeddings, the numeric vectors a language model computes for a text, plus logistic regression. With 400 examples it reached 85%, Jev with no examples 86%.
The verdict back then: without labels, Jev gives you a good start. With labels, try the baseline first. Part 2 checks whether that still holds after three weeks of development, with large open models, in German, and on tasks part 1 did not cover.
Two weeks after Jev, Cloudflare copies the interface
On September 30, Cloudflare released two models, Clef and Clef-flash. They use the same interface as Jev, read images and video as well as text, and are published as open weights under the Apache 2.0 license on Hugging Face. You can use them through Cloudflare’s API or download and run them yourself. (Cloudflare, Hugging Face)
A startup gives zero-shot classification a new name, and two weeks later one of the largest infrastructure providers offers the same thing, open and multimodal. Have decision models become interchangeable?

Things moved fast. Two days after the launch, the community was maintaining its own comparison, the “Jev Decision Index”. On September 29, OpenAI announced its own Decisions API, so far only as a limited preview. In early October, the model marketplace OpenRouter listed decision models as a category of their own. And on the morning of October 4, Clef was number 1 on the Hugging Face trending list, the main platform for open models. Laya, another open decision model, was number 2, CLM number 8. Popularity says little about fitness for the job, though: Laya and CLM score below 45% on my task. (Decision Index, The Decoder, OpenRouter)
A lot of marketing for a zero-shot classifier
Part 1 described Jev as a well-made product, and I stand by that. Price, speed and stability are real, and the technical documentation is remarkably candid, including about its own weaknesses. The vocabulary around the launch, on the other hand, consists mostly of new names for familiar things.
“System One” refers to Daniel Kahneman’s fast, intuitive thinking, which Kahneman describes as error-prone and overconfident. In AI research, fast models without step-by-step reasoning have been called that since 2019. “Noul” is a yes-or-no question with a probability, in other words binary classification. “Can’t hallucinate” only means that Jev picks from a fixed list. It can still be wrong, as the same website admits. And according to TypeSafe’s own primer, answers with a probability of 1.0 should always be right. In my test, 206 of 400 answers carried that value, and 11 of them were wrong. (TypeSafe, Primer)
On top of that came 40 million US dollars in seed funding at launch, led by the investor DCVC. For a startup whose first product is technically a zero-shot classifier, that is remarkable. Whether the models behind the names are interchangeable, however, is not settled by a term or a funding round. Only a test can show that. (DCVC)
The test: the same data as part 1, plus new questions
Every model got the same 400 papers and the same constructed test cases as in part 1, with exactly the same request as Jev. That way every number can be compared directly with part 1.
I tested 15 decision models from Cloudflare, Perplexity, Fastino, Liquid, Upstage and other vendors, through their APIs, on a laptop with 16 GB of memory, or on rented Nvidia GPUs (A40 and A100). OpenAI’s Decisions API is not included because it was not generally available at the time of the test.
Four tasks are new, ones part 1 did not cover:
- Changing options. Given the abstract of a paper, the model has to find the right title among five similar ones. Each case has different titles. This is a strength decision models claim for themselves: a classic classifier learns fixed categories and cannot handle options that change from case to case.
- Checking a claim. Is a sentence supported by a text? Sometimes the sentence comes from the text, sometimes from another text, sometimes a single number in it has been changed.
- Counting. How often does a certain event, such as a failed login, appear in a log? The answer is given on a scale from “never” to “six times or more”.
- German. How well do the models handle German texts? For this I used 600 research projects funded by the FWF, the Austrian Science Fund. Each funded project has one abstract in German and one in English. The models had to identify the research field. (FWF)
The rules are the same as in part 1: the protocol was fixed before each run, everything added later is marked, and a second model independently recomputed the results. All API calls and rented GPUs together cost about 11 US dollars.
Clef is as accurate as Jev and freely available
Clef gets 85.8% of the 400 papers right, Jev 86.3%. That is not a measurable difference. Clef-flash, the smaller model, is four points behind. So the large open models have caught up: in part 1 the best open model was six points behind Jev, although only small models ran there, on the laptop.
Like pplx-decider and Kev 4B by now, Clef has open weights. If you do not want to send your data to an API, you can run a model at this level yourself. That did not exist when Jev launched. Through Cloudflare’s own API, on the other hand, Clef costs 0.155 US dollars per 1,000 decisions, four times as much as Jev.
Clef is also better placed when it comes to confidence. Not a single answer carries a confidence of 100%, and its values are a good guide to which answers are likely to be wrong.
Through Cloudflare, however, Clef had two weaknesses: it missed the deciding sentence in long texts, and with ten papers per request it lost 20 points. So I ran Clef myself on a rented GPU, an Nvidia A100 with 80 GB of memory. On the 400 papers it gave the same answer as through Cloudflare in every single case. But both weaknesses were gone: with ten papers per request Clef did not lose a point, and in texts of up to 8,000 tokens it always found the deciding sentence. The model was never the problem. The problem was how Cloudflare serves it: long inputs are apparently truncated there without notice.
The code Cloudflare ships with the model has a limit of its own, though. It truncates every input to 16,384 tokens, without a warning. In the longest test, self-hosted Clef was therefore wrong almost one time in five. The limit can be raised with a single line. Set to 65,536 tokens, Clef found the sentence in all 28 cases, even in the longest texts. That costs time: for a text of 24,000 tokens, Clef needed 11.5 seconds on the A100 on average, compared with a third of a second for one paper.
Accuracy barely separates the models, price does

Six of the 15 models cannot be told apart from Jev on the 400 papers, among them Perplexity’s pplx-decider, Fastino’s GLiDE and Clef. pplx-decider even scores one point above Jev, but that is not a reliable difference either. The remaining nine are 4 to 47 points behind.
Is the task simply too easy, then? To some extent, yes. Already in part 1, the best systems were close to what the labels allow: on almost every paper that all of them got wrong, they agreed on the same category, a different one from the authors’. For practical purposes the result still means something. Many assignment tasks in a company are probably similarly simple, such as routing tickets to a department. There, what matters most is what comes after accuracy.

Price, for example. Among the equally accurate models, 1,000 decisions cost between 2 and 26 cents, a factor of up to 13. At 3.5 cents, Jev is at the low end.
In practice: Do not pick a decision model by its accuracy on a leaderboard. For simple classification, many are equally good. What matters is how the model behaves in production and what it costs.
Same interface, different behavior
Every model got exactly the same request. The answers look the same, but they do not mean the same thing.
The confidence value. Every answer contains a field called confidence. The name alone does not make it a confidence you can rely on. With some vendors it is the probability of the chosen answer, with others a score derived from it. If you build a threshold on it, such as “below 80%, a human checks”, you have to set it again every time you switch vendors.

Long texts. A long administrative text of up to 24,000 tokens, roughly 18,000 words, contained a single deciding sentence. Only Jev, pplx-decider, GLiDE and D1 found it at every length, plus self-hosted Clef once its text limit is raised. Other models simply reject long texts.
Many cases per request. If you send ten papers in one request instead of ten separate ones, some models keep their accuracy and others lose up to 50 points. The price flips as well: with Jev, batching gets cheaper, with most other vendors considerably more expensive.
No fitting category. Given a “none of these” option, almost all models recognize papers that fit no category. Some, however, also reject many papers that do fit, Solar Decide almost one in three.
In practice: With decision models, switching vendors means switching models, even if the code stays the same. Check beforehand what the fields mean, how long texts are handled and what a batched request costs.
How many cases a model can decide on its own
In production, the real question is rarely how accurate a model is overall. It is: which cases can I leave to it without review?
An example: a model with 86% accuracy gets roughly one case in seven wrong. If you let it decide alone only the cases it is most confident about and pass the rest to people, the error rate among the automatically decided cases goes down. I measured what share of cases a model can take over this way if at most one in twenty of them, 5%, may be wrong.

Here, equally accurate models differ widely. GLiDE, Clef and pplx-decider could each take over a little more than half of the cases. Jev, not a single one. The reason is the finding from part 1: Jev gives half of all answers a confidence of 100%, and a little over 5% of that half is wrong. That half cannot be sorted any further. With a target of 5%, nothing is left; at 5.5% it would already be half. All or nothing.

These numbers are an estimate with a lot of uncertainty. They show that equally accurate models differ precisely in the property that matters for automation. They do not predict how much a model can take over on your data.
In fairness: if you state the probability explicitly, for example with an urn containing seven red and three blue balls, Jev reproduces the probability more accurately than any other decision model. Jev’s problem is not probabilities as such, but judging its own mistakes.
In practice: Before deployment, measure on a few hundred of your own cases how many the model can decide alone at the error rate you can tolerate. That number determines how much work you save, and it can vary several-fold between equally accurate models.
Changing options: the real strength, rarely needed
A correction to part 1 belongs here. Part 1 said a Choice question is classic classification. That is true as long as the options are fixed. A decision model, however, can receive options that change from case to case, and a classic classifier cannot do that. This is the core of what decision models can do beyond classification.
In the test it worked well. When picking the right title among five similar ones, almost all decision models scored between 97 and 99%. But a simple comparison of embeddings, with no training at all, also reached 97%. So decision models can handle changing options, but in this test they were not the better tool for it. Whether that changes with harder options, I did not measure.
How often do you need this in practice? In my experience, rarely, and especially not in the advertised cases such as routing emails or tickets. Those have fixed categories. Changing options mostly come up when an AI agent has to pick one tool out of many, or when records are matched against each other. Even there, the vendors themselves usually solve it in two steps: a simple search picks a short list, and the model then checks each candidate one at a time. That is also how TypeSafe’s own guides describe it. (TypeSafe docs, Anthropic)
The other two new tasks gave a mixed picture. Checking claims against a text went well for most models; only small models stumbled over changed numbers. Counting events in a log, by contrast, was hard for all of them. Even the best got just under two thirds of the levels right.
German costs hardly any accuracy
I ran the language test only with ten hosted models. On the 600 FWF projects, nine of them stayed within one percentage point between German and English. So the models handle German texts well. German texts are longer, though: they need a fifth to a third more tokens, and with per-token pricing they cost correspondingly more.
Two caveats. The two abstracts of a project are not translations; applicants write them separately. The test therefore compares the languages as applicants write them, not pure translations. And the texts have been public for some time, so the models may already know them. That does not skew the language comparison, because every project appears in both languages, but it does affect the absolute accuracy.
With labels, a simple baseline is enough
In part 1, a simple baseline of embeddings plus logistic regression came close to Jev with 400 labeled examples. The training examples there were 2,806 papers from 2024, separate from the 400 test papers from September 2026, and the learning curve drew samples of different sizes from them. That raises a question that has stayed with me since part 1: some decision models, Laya for example, are explicitly offered as a starting point for fine-tuning. Why use a decision model at all if you have to train it anyway? You might as well start with an ordinary model.
I tested this after the fact. Laya and ModernBERT, the model Laya is built on, were fine-tuned on the same labels. ModernBERT got an ordinary classification head for this.

With 400 labels both reached just under 80%, the baseline 85%. With all 2,806 labels the ordinary model reached 87%, Laya 84%. Pretraining as a decision model brought no advantage.
The way there is more instructive than the result. Just to make the fine-tuning run stably, it took a whole pipeline of the usual machine learning tools: searching for a suitable learning rate, holding out a separate validation set, choosing the best training length, and repeating every setting several times with different random seeds. That came to 42 training runs and about five hours on a rented Nvidia A40, only to end up just under 80%. The baseline trains on ready-made embeddings in seconds, without a GPU or a validation set, and comes out five points ahead.
If you decide to fine-tune, you should know this. On top of the labels come experiments, compute and expertise. For large language models the effort is likely to be even greater, although I did not measure that.
In practice: If you have labels and fixed categories, train the simple baseline first. Fine-tuning a decision model did not pay off in this test.
Which model I recommend for which situation
The recommendations apply to what I measured: assigning texts to fixed categories in English and German, as of October 2026. Hosted models can change at any time, and how many cases a model can decide alone is an estimate.

- You have labels and fixed categories: Try the simple baseline of embeddings plus logistic regression first. In this test, at 400 labels, it was more accurate than any fine-tuned model and the best at judging its own uncertainty.
- No labels, hosted, price matters: pplx-decider from Perplexity. As accurate as Jev, among the cheapest at 2.2 cents per 1,000 decisions, and it reads long texts in full. Batched requests, however, cost several times as much.
- Many cases should run without a human: GLiDE from Fastino, or pplx-decider. In my estimate both could take over a little more than half of the cases on their own, but GLiDE is the most expensive model. Check the share on your own data.
- High volume, batched requests, a human reviews: Jev remains a good choice. It is the most stable when categories are reworded, and batching makes it cheaper. In my test, however, Jev could not take over any cases entirely without review at an error rate of at most 5%.
- The data must stay in-house: Clef from Cloudflare with open weights, on your own hardware. Raise the text limit in the code that ships with it, or Clef will truncate long texts without warning. For smaller GPUs, Decision 2.0 Nox 4B is an alternative, though it rejects very long texts. pplx-decider and Kev 4B are also available with open weights, but I only tested both through their APIs.
- Someone has to be able to follow the decision: a language model with a rationale, as in part 1. Decision models do not explain themselves.
Less suited to this kind of task were Solar Decide, which rejected many fitting cases, and Kev 4B on long or batched texts. CLM and Laya are built for other purposes: choosing agent actions, and serving as a base for fine-tuning.
What this shows and what it does not
The test mainly covers one kind of task: assigning texts to eight fixed categories. The new tasks are constructed, and some are close to being solved by every model. I did not test decisions with rules, exceptions or several related questions, nor images, even though Clef can read them. Everything in this part is post hoc relative to part 1. Every step is logged, and steps I decided only after seeing intermediate results are marked in the repository.
Hosted models often have no public version number and can change at any time. Local models ran on a laptop with 16 GB of memory, Nox 4B and the fine-tuning on an Nvidia A40, self-hosted Clef on an Nvidia A100. Response times are not comparable across these setups. The Hugging Face numbers are a snapshot from October 4.
How this text was made
The experiments ran from October 2 to 4, 2026: the local models overnight on the laptop, some models and the fine-tuning on rented GPUs that the agents set up themselves and deleted afterwards. As in part 1, AI agents in Claude Code did the work: Claude Opus 5.5 planned, wrote code, ran the experiments, researched and drafted the text. Codex with GPT-6 Sol independently reviewed the report, the key numbers and the structure of this text. I set the question, made every decision about models, data and budget, checked intermediate results and revised the text. This English version is an adaptation of the German original by a Claude agent, reviewed by Codex.
Do you even need a decision model?
For the task I measured, assigning texts to fixed categories, the answer is: if you have labels, probably not. Try the simple baseline first. In this test it beat every fine-tuned decision model and trained in seconds.
If you do not have labels, a decision model can be a good start. Jev still is one, but no longer the only good choice. On accuracy, many are now level, including models with open weights. Choose by what comes after: how many cases the model can decide alone on your data, how it handles long texts and batched requests, and what it costs.
Decision models only play to their real strength when the options change from case to case, for example when an agent picks from many tools. That happens less often than the marketing suggests, and even there, in my test, a simple comparison of embeddings was almost as good.
So the conclusion of part 1 still holds, with one addition: if you have no labels yet, you now have a choice. Accuracy hardly decides it anymore. If the model is to decide cases on its own, what matters most is how well it judges its own mistakes.
The full test, all data, every model response and the analysis are in the public companion repository: decision-model-audit. It also lists every deviation from the original plan, so you can check the results rather than take my word for them.
Sources
Part 1 and companion repository: Is Jev Actually Any Good? A Field Test on Fresh Data; decision-model-audit, report results/REPORT-part2.md
TypeSafe and Jev: Announcement; Website; System One in the docs; Machine learning primer; Docs overview; Funding, DCVC
New decision models: Cloudflare Clef, Clef on Hugging Face; Perplexity pplx-decider; Fastino GLiDE; Liquid D1 on OpenRouter; Decision 2.0 Nox 4B; APUS-OpenJev; Strands Decider; CLM; Laya; ModernBERT-large
Developments and context: Jev Decision Index; The Decoder on OpenAI’s Decisions API; Anthropic on tool search for agents
Data: FWF Open API; arXiv as in part 1
Facing a complex problem? Before any code, a conversation. No pitch, no commitment.
Book a call