Skip to content
Is Jev Actually Any Good? A Field Test on Fresh Data
ai-evaluationclassification

Is Jev Actually Any Good? A Field Test on Fresh Data

TypeSafe's Jev tested against small LLMs, open alternatives and the old recipe of embeddings plus labels, on fresh data from September 2026.

TL;DR. A technical post that goes deep into machine learning and evaluates Jev, mainly for people who build or select models themselves.

  • Against small LLMs: Jev is not measurably better or worse, but it is cheaper and faster. The marketing numbers come from vendor tests against far more expensive models.
  • Against open alternatives: Jev is clearly ahead.
  • With labels: A simple classifier on modern embeddings almost catches up with Jev after a few hundred examples, and it is better calibrated.
  • Probabilities: Usable, but often saturated and dependent on how you ask.
  • Verdict: A good zero-shot classifier for the cold start, not a new principle.

Transparency note: This text was written with AI, and AI agents ran the experiments. I planned, supervised and checked both. How that worked is described at the end. Code, data and every model response are in the linked repository.

On September 15, 2026, TypeSafe presented its first model. It is called Jev, and according to the company it is a “System One” model: it does not write text, it decides. You give it a text and a question with fixed answer options, and it returns the answer along with a probability. Two numbers came with it and spread quickly: 193.6 times faster and 444.6 times cheaper than a large language model. (TypeSafe, Tom’s Hardware)

The announcement traveled fast. I still raised an eyebrow, and I was not the only one. (KDnuggets) A model that picks from a list of categories without any training data of its own, and attaches a probability, has had a name for a long time: a zero-shot classifier. A look around LinkedIn might still suggest that classification was invented last week. The announcement works as a quiet litmus test. Anyone who has worked with machine learning recognizes the underlying principle at once.

So zero-shot classification has been given a new name. In our industry, the renaming is often the most labor-intensive part of the innovation.

That alone would not justify a long post. New names for old ideas do not automatically make bad products, and TypeSafe’s numbers can be checked. Routing decisions are part of my daily work. I am responsible for a routing pipeline in customer support that I scaled to 5 million tokens per minute, and some time ago I published fse, a library for embeddings. So I checked the numbers, with a practical question in mind: what do you buy with a zero-shot model, and what changes once you have labeled examples?

#How classification works

Classification assigns one of several categories to an input. Is this email a cancellation, a billing question or an upgrade? Does this paper belong to robotics or to computer vision? A model usually outputs a probability per category. A threshold lets you decide when the system decides on its own and when a human takes a look.

There are two basic routes. In supervised learning, you train a model on examples whose correct category you know, in other words on labels. This ranges from TF-IDF, which weights words by how rare they are, combined with a logistic regression, all the way to embeddings. Those are numeric vectors that a language model computes for a text. A simple classifier then learns on top of them. In the zero-shot approach, there are no examples. The model only gets a description of the categories and has to infer the answer from what it already knows. Incidentally, this has nothing to do with unsupervised learning, where a method looks for structure in data without any labels at all, as in clustering. The categories are given. Only the examples are missing.

Neural networks producing decisions is the normal case, by the way. The softmax layer at the end of a classifier is their usual output form. Strictly speaking, a large language model also classifies at every single token, a piece of a word, just over a vocabulary of a hundred thousand entries.

That leaves the question of whether the probabilities are right. This is called calibration. A model is well calibrated when its confidence keeps its promise: if it says 90% on each of a hundred decisions, roughly 90 of them should be correct, not 70 and not 99. Only then can you build a threshold on that number, such as “below 80% confidence, a human checks”. That modern neural networks are systematically overconfident has been well documented since 2017 at the latest. (Guo et al.)

#What Jev is and what is supposed to be new about it

Jev receives a “state”, TypeSafe’s somewhat grander word for the input, usually a text, plus one or more typed questions. It returns the answer and probabilities, but no justification. Several questions about the same text can be asked in a single call. TypeSafe has not published how the model is built internally or how large it is. (TypeSafe Docs)

The three question types are the core of the interface, and they come up again in every comparison:

  • Choice picks one of up to 255 options that you define in the call, and returns a probability distribution over all options. This is classic multi-class classification: which department gets this ticket?
  • Noul answers a yes-or-no question with the probability of yes. This is binary classification: does this email contain a cancellation?
  • Score places something on a scale with fixed levels and returns a probability per level. This is ordinal classification: how urgent is this case, from 0 to 4?

None of these three types is new. What is new is that a single model serves all three through one shared interface.

The company’s reasoning is more interesting than the product sheet. After pretraining, large language models usually go through a procedure called RLHF, reinforcement learning from human feedback: people rate answers, and the model learns to give the answers people prefer. Diogo Almeida, co-author of InstructGPT, one of the papers that made RLHF widely known, and founder of TypeSafe, argues that this is exactly what makes language models eager to please and overconfident. That is fine for an assistant, he says, but not for automated decisions. TypeSafe therefore trains with its own procedure, RLCD, so that the probabilities match reality. The abbreviation stands for Reinforcement Learning for Calibrated Decisions. There is no technical paper on it so far. (Talk by Diogo Almeida, AI Engineer World’s Fair, MarkTechPost)

Is this the right approach? The diagnosis has a solid core. OpenAI already showed in its GPT-4 report that the pretrained model was well calibrated and that calibration got worse after post-training. (OpenAI, GPT-4 Technical Report) The cure is a bet. Calibration can also be added after the fact, with methods such as temperature scaling, and has been for years. Whether it needs its own training procedure only shows in the results. More on that below.

Either way, the idea of choosing from categories without training data is considerably older than Jev.

Timeline of zero-shot classification, from dataless classification in 2008 through NLI-based classifiers in 2019 and GLiNER in 2023 to Jev in September 2026.

In 2008, Chang and colleagues classified texts without labeled data. In 2019, Yin and colleagues framed classification as an inference task, natural language inference (NLI): does the sentence “This is about sports” follow from the text? The zero-shot pipelines you find in many tutorials are based on this idea. GLiNER followed in 2023. Its successors now sit inside one of Jev’s open competitors. Structured outputs, meaning answers guaranteed to follow a given schema, are now offered by the major providers, including OpenAI, Anthropic and Google. (Chang et al., Yin et al., GLiNER, OpenAI)

The task itself, then, is not what is new about Jev. What can be new is the training objective, the price, the speed and the promise that the probabilities are right. That can be measured.

Shortly after the launch, the first open alternatives appeared: Eikos, Kev, SemIf, plus Fastino’s GLiNER2.5-Decide. Others, such as GLiClass or Laya, already existed. Two weeks in, there were several alternatives whose developers openly challenge Jev. Whether the lead holds is what the test shows.

#The test: fresh data, someone else’s labels

The biggest problem with comparisons like this is the test data. Public benchmarks such as AG News have demonstrably ended up in the training data of some large language models. A test whose answers the model already knows measures nothing. (Golchin and Surdeanu) Labeling the data myself was out of the question, because my labels would then have shaped the result. Besides, I have labeled more than 15,000 tweets in my life. Once is enough. If you have done it, you know.

So I used data whose labels were assigned by an institution before any model could see them. The main task is 400 arXiv papers from September 2026, 50 from each of eight computer science categories: language, computer vision, security, robotics, databases, software engineering, human-computer interaction and distributed computing. The label is the primary category that the authors chose and the moderators confirmed. An example, shortened:

Bridge3D: Enabling Vision-Language-Action Models to See and Act in 3D

Vision-Language-Action (VLA) models have demonstrated remarkable generalization
in robotic manipulation via large-scale multimodal pretraining. However, ...

Label: cs.RO (Robotics)

On top of that come 358 documents from the Federal Register, the official journal of the US federal agencies, labeled with one of six issuing agencies. I masked the agency names in the text:

Airworthiness Directives; Bombardier, Inc., Airplanes

The [AGENCY] proposes to adopt a new airworthiness directive (AD) for all
Bombardier, Inc., Model BD-100-1A10 airplanes. ...

Label: FAA

Both are Choice questions. The set is rounded off by four kinds of constructed test cases whose correct answer is fixed by how they are built. Two of them are Noul questions: an urn with seven red and three blue balls and the question of whether the ball drawn is red, and a long administrative text with the question of whether it contains a cancellation anywhere. I did not test Score questions. More on that below.

Two limitations belong right here. The data is newer than every disclosed training cutoff of the language models tested. Jev, Gemini and the open models do not disclose their cutoffs. And a paper’s category is the authors’ choice between overlapping fields. A paper on fuzzing, automated testing with random inputs, can sit under security or under software engineering. For 16 of the 400 papers, all five API models tested agreed on a different category than the authors. So 100% is not realistic on this task.

In total, 23 systems ran. The protocol was fixed before the first call. Every deviation from it is documented in the repository, and everything I added only after the first results is marked as post hoc. The bill for all API calls came to about $5.30.

The systems differ a lot, so here are the most important ones at a glance:

SystemTypeAccessContextInputsProbabilityJustificationCost per 1,000
Jev 1.13Decision modelAPI only32,000 tokensTextyes, from the modelno$0.035
GPT-6 LunaLanguage modelAPI1M tokensText, image, PDFself-estimated onlyyes$0.065
Gemini 3.5 Flash-LiteLanguage modelAPI1M tokensText, image, PDF, audio, videoself-estimated onlyyes$0.21
Claude Haiku 4.5Language modelAPI200,000 tokensText, image, PDFself-estimated onlyyes$0.89
Claude Sonnet 5Language modelAPI1M tokensText, image, PDFself-estimated onlyyes$2.48
GLiNER2.5-DecideOpen encoderlocal512 tokens trained, chunking possibleTextyesnoown hardware
LayaOpen encoderlocal512 tokensTextyesnoown hardware
DeBERTa-v3 NLIOpen encoderlocal512 tokensTextyesnoown hardware
Embeddings + logistic regressionSupervised, needs labelsAPI or local32,000 tokensTextyesno$0.003

The comparison covers what all of them can do: same text, same categories, one answer. That a language model can also classify images and PDFs and write justifications, while Jev only understands text, appears in the table as a capability, not as a score. It is still apples and oranges.

#How the models were asked and how they answered

Every system has its own interface. All of them received the same text and the same category descriptions, each in its own format. Jev gets the text as the state and a Choice question with the categories as criteria:

{
  "state": "Bridge3D: Enabling Vision-Language-Action Models to See and Act in 3D ...",
  "questions": {
    "q": {
      "type": "choice",
      "instructions": "Which arXiv category is the primary category of this computer science paper?",
      "criteria": {
        "cs.CL": "Computation and Language: natural language processing, ...",
        "cs.RO": "Robotics: robot control, manipulation, locomotion, ..."
      }
    }
  }
}

What comes back is the choice, a confidence and a probability per category, for example "choice": "cs.RO", "confidence": 1.

The language models get the same question and the same categories as a system instruction, the text as a message, and a fixed response schema via structured outputs. The label field may only contain one of the eight categories:

{
  "label": {
    "type": "string",
    "enum": ["cs.CL", "cs.CV", "cs.CR", "cs.RO", "cs.DB", "cs.SE", "cs.HC", "cs.DC"]
  },
  "confidence": { "type": "number" }
}

The answer then looks like this: {"label": "cs.RO", "confidence": 0.99}. The confidence is a number the model writes down itself, not a probability from its internals. The providers do not expose those via OpenRouter.

The NLI model checks for each category whether the sentence “This text is about Robotics: robot control, manipulation, …” follows from the text, and compares the scores. GLiNER, Laya and the other decision models each use their maker’s library and return scores per category.

#Against small and mid-sized language models, no difference is measurable

On the 400 arXiv papers, Jev matches the authors’ category in 86.3% of cases. GPT-6 Luna reaches 88.8%, Gemini 3.5 Flash-Lite 88.0, Claude Haiku 4.5 87.3 and Claude Sonnet 5 89.8%. All four language models are therefore slightly ahead of Jev on the point estimate. After correcting for multiple testing, none of these differences is statistically robust at 400 cases. That does not mean the systems are equally good. It means this test cannot tell them apart.

Accuracy of all systems on 400 arXiv papers with 95% confidence intervals. Jev sits among the small and mid-sized language models, while the open zero-shot models are clearly below.

Part of the reason lies in the task itself. It is probably too easy for today’s language models, and the best systems are close to what the labels can support at all. All five API models are wrong together on only 18 papers, or 4.5%, and on 16 of those they agree on the same other category. There is little room left between good and very good. A harder task might separate the systems more clearly. On the Federal Register documents, the ceiling was even closer: Jev and all four language models landed between 99.2% and 99.7%, with no measurable difference. The open models reached 81% to 98% there.

A repeat run shows how little a single run says. I sent 98 papers to the same models a second time. Jev and Haiku answered identically in 99% of cases. GPT-6 Luna changed its mind on one paper in twenty. Its accuracy on these cases dropped from 89.8 to 85.7%. At least for GPT-6 Luna, then, differences of a few points are within what a second run shifts on its own.

On price, Jev is ahead. Via OpenRouter, 1,000 decisions cost $0.035 with Jev, $0.065 with GPT-6 Luna, $0.21 with Gemini, $0.89 with Haiku and $2.48 with Sonnet. Jev’s median response time was 0.47 seconds, the language models’ 0.77 to 2.14 seconds, and GPT-6 Luna’s 1.14 seconds. That makes Jev about 2.4 times as fast as GPT-6 Luna and 1.9 times as cheap. If you send ten papers in one request, costs for Jev and GPT-6 Luna drop by about a quarter. Per decision, Jev is then about 3.3 times faster and 1.8 times cheaper than GPT-6 Luna.

Accuracy against cost per 1,000 decisions on a logarithmic axis. Jev is the cheapest API system, and GPT-6 Luna costs just under twice as much at similar accuracy.

That is a real advantage, but not a factor of 400. TypeSafe’s factors come from its own workflow tests, in which Jev runs against a range of language models, from Claude Haiku 4.5 to Claude Opus 5. GPT-6 Astra and Claude Fable 5.1 only supplied the reference answers there. The two factors cannot be recomputed directly from the published tables, but their order of magnitude fits the most expensive models in the comparison, such as Opus 5. Next to a model like that, a lot of things are 200 times faster. GPT-6 Luna in TypeSafe’s own table is more telling: there, Jev is about 8 times cheaper and roughly 30 times faster, probably because every decision in those workflows consists of many individual questions. In my test, with one question per paper, the factors are 1.9 and 2.4. If you already make these decisions with a small model, you are more likely to gain a factor of two to three. (TypeSafe workflow evals, TypeSafe)

Where Jev genuinely stands out is stability. If you reverse the order of the categories or describe them differently, 97.5 to 99% of its answers stay the same. For GPT-6 Luna it is 94 to 95%, for Gemini 92 to 93%. For a system that has to make the same decision thousands of times in the background, that is a property you notice in production.

#An independent replication with other question types

Almost at the same time, Stefan Rossmeier published his own comparison of Jev and GPT-6 Luna, with a completely different setup: 5,760 synthetic decisions from 16 business areas, spread evenly across Choice, Noul and Score, with labels from fixed rules. There, Jev is 1.4 points ahead of GPT-6 Luna overall, and on Choice the two are practically level. Jev’s lead comes mainly from the Noul questions. GPT-6 Luna says no conspicuously often there and therefore misses roughly one in four cases where the correct answer would be yes. Jev is balanced in both directions. On Score, all systems hit only around 40% of levels exactly, and Jev’s calibration there is clearly worse than on Choice, with an error of 0.31. Rossmeier also measures Jev at about three times as fast and about 1.7 times as cheap as GPT-6 Luna. (decision-model-evals)

One detail puts the 5,760 into perspective: the cases are generated from 160 base scenarios, each combined with twelve fixed context sentences. If you evaluate per scenario instead of per case, the 95% interval for Jev’s lead runs from minus 0.3 to plus 3.0 points. That recalculation is mine, based on his published raw data. Both tests thus reach the same conclusion from opposite directions: between Jev and a small language model, there is no robust difference in accuracy, but there is one in price and speed.

In practice: If you already use a small language model with structured outputs for simple classification, Jev will not reliably give you higher accuracy. What you gain is price, speed and consistency. For yes-or-no questions, it is worth checking whether your language model favors one side. Whether an additional vendor is worth it depends on your volume. At a million decisions per month, the difference against GPT-6 Luna in my test is around $30.

#The open alternatives fall clearly behind

Against the open models I could test on a laptop with 16 GB of memory, Jev is clearly ahead. GLiNER2.5-Decide, Fastino’s direct answer to Jev, gets 73.5%, 12.8 points less. The NLI model by Moritz Laurer, a classic zero-shot classifier from before Jev, reaches 66.8%, GLiClass 64.3 and Laya 43.5%. I later added Eikos-4B at 80.5 and SemIf at 77.5%. All of these gaps are statistically robust.

This seems to contradict the LangWatch comparison titled “Open models caught up with Jev”. There, open models with up to 28 billion parameters are level on eight of eleven tasks. The two results do not rule each other out. LangWatch measures on public datasets that models may have seen during training, and I could not test the large models with 26 to 27 billion parameters on the laptop. For models that run on a laptop, however, the claim does not hold on fresh data. (LangWatch)

A side finding shows how much the result depends on the way you query. Qwen3.5-4B is a small open language model that I queried on the laptop without structured outputs. Instead of letting it write an answer, I listed the categories as letters and read off directly how likely the model was to write A, B, C and so on next:

Which arXiv category is the primary category of this computer science paper?

Options:
A) cs.CL: Computation and Language: natural language processing, ...
B) cs.CV: Computer Vision and Pattern Recognition: ...
...
H) cs.DC: Distributed, Parallel, and Cluster Computing: ...

Answer with the letter of the correct option only.

That way it reached 53.3%, mostly because it picked A remarkably often, regardless of what the text said. SemIf uses the same model but asks with its developer’s prompt and readout method, and reaches 77.5%. Same model, 24 points apart, purely because of the prompt and the readout method.

In practice: An open model on your own hardware is attractive for data protection and cost, but the numbers from launch posts and leaderboards do not automatically carry over to your data. Test every model with your own query on your own examples before you commit.

#What Jev’s probabilities are worth

Jev’s real promise is not the answers but the probabilities that come with them. They are supposed to let you decide confident cases automatically and pass uncertain ones to a human.

On average, they are reasonably well calibrated. The expected calibration error (ECE) summarizes how far the stated confidence deviates, on average, from the actual accuracy. Zero would be perfect. Jev comes in at 0.068, on average just under seven percentage points off, which puts it on par with the language models. The problem lies in the distribution. For 51.5% of all answers, Jev states a confidence of exactly 1.0. Within that half, there is no way left to sort which cases a human should check first. Rossmeier gets 48.7% on his data, so the pattern is not a quirk of my sample.

Error rate among the automatically decided cases as a function of the automated share. Jev's curve is flat up to 51.5%, because half of its answers carry a confidence of 1.0.

It becomes clearer when you ask the same question differently. Instead of one Choice question with eight categories, I asked Jev eight Noul questions, one per category: “Is the correct category cs.RO?” The chosen category almost always stayed the same. The calibration error, however, rose from 0.068 to 0.251.

There is a second finding. Eight separate yes-or-no questions are like eight separate classifiers, similar to eight individual logistic regressions. For any single case, their yes probabilities do not have to add up to exactly 1. But because every paper has exactly one primary category, well-calibrated answers should add up to 1 on average. For Jev, the median sum was 1.25, and for 82% of the papers it fell outside the range of 0.9 to 1.1. So on these yes-or-no questions, Jev systematically says yes too often. On average, the probabilities add up to 125%, a figure otherwise reserved for motivational speakers. An independent audit by a UCLA researcher found the same pattern on a different dataset. (Zenodo, alexmolas.com)

Jev's calibration by question form. The calibration error is 0.068 as a Choice question and 0.251 as Noul questions, shown alongside the distribution of summed yes probabilities with a median of 1.25.

Another criticism, on the other hand, I could not confirm. According to alexmolas.com, Jev gave a probability of 0.92 for the toss of a fair coin. On Noul questions with an explicitly stated probability, such as the urn with seven red and three blue balls, Jev was off by only 0.027 on average. And on the question of whether any category fits at all, Jev is strong. With an explicit “none of these” option, it recognized in 95 of 100 papers from physics, mathematics, biology and economics that nothing fits, and wrongly said “none” for only 1% of the matching computer science papers. Without that option, every answer is necessarily wrong. In that setting, Jev gave 11% of its answers a confidence of at least 0.9, GPT-6 Luna 53%.

In practice: Use Jev’s probabilities for sorting, not as absolute truth. Ask a question with several categories as a Choice question, not as a series of Noul questions, and offer a “none of these” option. According to Rossmeier’s data, Score questions call for particular caution. Before you set a threshold, check on a few hundred of your own cases how many errors actually happen above it, and recalibrate if needed.

#With labels, a simple classifier catches up

Up to this point the comparison group was clear: zero-shot against zero-shot. The more interesting question is what happens when you have examples. For that, I later tested a well-known recipe from supervised learning. Each text is turned into a vector by an embedding model, and a logistic regression learns the categories on top of it. The training data was 2,806 arXiv abstracts from 2024, and the test used the same 400 papers from September 2026. So almost two years separate training and test.

With all 2,806 examples, Qwen3-Embedding-8B with logistic regression reaches 86.3%, the same score as Jev. The calibration error is 0.029 instead of 0.068, handing uncertain cases to humans works best of all systems on the point estimate, and via OpenRouter the embeddings cost about one twelfth of what Jev costs per decision. The embedding model is open and could also run locally. The differences in accuracy are not robust, and I did not run a separate test for calibration. A TF-IDF model, whose weighting dates back to the 1970s, reaches 82 to 83% with the same training data.

What matters is how many examples it takes.

Learning curve on 400 arXiv papers. With 20 examples per category, the embedding classifier reaches 82.4%, with 50 per category 85.0%, while Jev reaches 86.3% without any examples.

With 20 examples per category, 160 in total, the embedding classifier reaches 82.4% on average over ten draws. With 50 per category it is 85.0%, 1.3 points below Jev, and at that point it is already ahead when it comes to handing cases to humans. TF-IDF needs about ten times as many examples for the same accuracy. The modern embeddings make the difference.

To be fair, two limitations belong here. My labels came free, because arXiv supplies them. In a real project you have to collect, check and maintain them, and I did not measure those costs. And the learning curve was added post hoc, on a single Choice task with eight categories.

A hint in the same direction comes from TextCortex. Their model Raya is a Laya model, fine-tuned for six minutes on a single routing task. According to the company, it is 3 to 4 points behind Jev on that task and just under 10 points ahead on a different question form. (Raya)

So Jev’s strongest case is the cold start: a decision for which no suitable labels exist yet. As soon as a few hundred examples exist, a recipe that appears in no launch announcement plays in the same league.

In practice: Before you buy a zero-shot model, look for labels that already exist: ticket categories, approvals, forwarded tickets, metadata. A few hundred of them are often enough for a classifier of your own that is cheaper, better calibrated and independent of any vendor. Jev then works as a starting point until that data is available.

#Justifications and long texts

What Jev fundamentally cannot do is justify a decision. A language model can output text next to the label that a human can use to check or challenge the decision. GPT-6 Luna and Gemini quoted verbatim from the paper in over 90% of cases, while Haiku did so only rarely. For decisions that someone has to be able to follow, that is a real difference.

The justification is not free. When I instructed the language models to first write a justification with quotes and then choose the category, their accuracy dropped by 4.5 to 7 points. Costs rose to 1.5 to 1.8 times as much. This applies to exactly this kind of instruction, not to justifications in general, and I did not measure whether the justifications actually help humans check the decision. Unfortunately, a fluent justification for a wrong answer is just as convincing as one for a right answer.

Accuracy of the three language models without and with a justification written first. GPT-6 Luna drops from 88.8 to 81.8%, Gemini and Haiku by 4.5 points each, at 1.5 to 1.8 times the cost.

Long texts also split the field. In a Noul test, the decisive sentence, a cancellation, was placed somewhere in up to 24,000 tokens of administrative text. Jev and GPT-6 Luna found it at every length. The fixed-window encoders, meaning Laya, the NLI model and GLiClass, only read the beginning and dropped to around 46 to 50% on long texts, which is chance level between yes and no. Without splitting the text into sections, the laptop did not have enough memory for GLiNER. With the documented chunking, it reached 86%. The sentence did stand out stylistically, though. This was a search test, not a test of understanding long documents.

Accuracy by text length from 500 to 24,000 tokens. Jev and GPT-6 Luna stay at 100%, while fixed-window encoders drop to around 50%.

In practice: If someone needs to be able to follow or challenge a decision, the language model remains the better choice. Account for the lower accuracy and higher cost, and when in doubt, have the justification written after the decision rather than before. For long documents, first check how much text a model actually sees.

#Which tool for which decision

The results lead to a set of questions you can work through in order.

Decision tree covering whether labeled examples exist, how many, and whether justifications or long inputs are needed, with the fitting tool for each case.

  1. First check whether labels already exist. They are often hidden in tickets, approvals or metadata that someone has been maintaining anyway. With a few hundred examples, a classifier on embeddings is worth trying before you buy a zero-shot model.
  2. Use a decision model like Jev for the cold start. High volume, a fixed question, no examples: this is where Jev was fast, cheap and stable in my test.
  3. Recalibrate the probabilities on your own data. With Jev, they depend on the question type, and half of the answers carry a confidence of 1.0. Without checking, they are no basis for thresholds.
  4. Reach for the language model when someone has to follow the decision. The same goes for long inputs. The language models’ capabilities speak for them on multilingual or multimodal inputs, but I did not test that.
  5. Avoid fixed-window encoders for long documents unless you split the text into sections. Whatever lies beyond the window does not exist for them.

#What this shows and what it does not

The test mainly covers one kind of task: Choice questions that assign English texts to six to eight categories. I only tested Noul with constructed test cases, and Score not at all. Rossmeier’s replication partly fills this gap, though with synthetic data and labels from fixed rules. I did not test decisions involving exceptions or calculation steps, nor other languages, nor tasks with very many options, where according to the Laya developers Jev is stronger. The whole thing is a snapshot from September 2026. Models behind an API can change at any time.

I did not re-check the labels of the arXiv papers, and the constructed test cases are artificial. Local models were limited by 16 GB of memory, which is why the large open models are missing. Apart from the repeat run with 98 papers, every condition ran once. I measured speed and cost one request at a time, not under parallel load. Many analyses, including the embedding baselines and the learning curve, were added only after the first results, and they are marked accordingly in the repository.

#How this text was made

The research, experiments, analysis and this draft came together in about two days, on September 27 and 28, 2026. The local models ran overnight on the laptop. The work was done by AI agents in Claude Code: Claude Opus 5.5 planned, wrote code, ran the experiments and drafted the text, and further Claude agents researched and cross-checked sources. As a second model, Codex with GPT-6 Sol independently reviewed the methodology, the results and the structure. Gemini transcribed the TypeSafe founder’s talk. I set the question, made every decision about data, models and budget, checked intermediate results and revised the text. This English version is an adaptation of the German original by a Claude agent, reviewed by Codex and checked by me.

#Back to the new name

Jev is not a new principle. It is a zero-shot classifier with a clean interface, a good price and remarkably stable answers. In my test, on 400 cases it could not be told apart from small and mid-sized language models, and it clearly beat the open alternatives tested. Its probabilities need checking on your own data. And if you already have labels, a well-known method and a current embedding model get you almost the same result.

Reading Jev as a leap in quality confuses price and speed relative to the most expensive models with a leap in substance. That is a shame, because the product is well made for what it is meant to do.

If you do not have labels yet, Jev gives you a good start. If you do, try the old recipe first.

The full test, all data, every model response and the analysis are in the public companion repository: decision-model-audit. It also lists every deviation from the original plan, so you can check the results rather than take my word for them.

One last question for you: how many labeled examples for your most important routing decision are sitting somewhere in your systems without anyone ever having thought of them as training data?


#Sources

TypeSafe and Jev: Announcement; Documentation; Workflow evals; Talk by Diogo Almeida, “I made ChatGPT, now I’m building what’s next”; MarkTechPost; Tom’s Hardware

Independent tests and assessments of Jev: KDnuggets; decision-model-evals, Stefan Rossmeier; Calibration audit, UCLA (Zenodo); alexmolas.com; jev-calibration-audit; LangWatch

Open alternatives: GLiNER2.5-Decide; Laya; Raya

Classification, zero-shot and calibration: Spärck Jones 1972, IDF; Chang et al. 2008; Yin et al. 2019; Zaratiana et al. 2023, GLiNER; Guo et al. 2017; OpenAI, GPT-4 Technical Report; OpenAI, Structured Outputs

Data contamination: Golchin and Surdeanu 2023

Companion repository: decision-model-audit

Dr. Oliver Borchers
Dr. Oliver Borchers
AI & Automation Advisor

Facing a complex problem? Before any code, a conversation. No pitch, no commitment.

Book a call