The Rise of Decision Models
In this blog
- Why decision models could matter
- How is Jev built?
- So, what are people actually doing with Jev?
- But, there are limitations
- Wait, things just got more interesting
- Kev: More control and more choices
- Strands Decider 2B: Small enough to live with the agent
- Clef: Decision models as infrastructure
- So how do they compare?
- Where I see it getting interesting
- Just the beginning
- Key Sources
- Download
Over the last couple of weeks, a new type of AI model has been getting a lot of attention: the decision model.
It started with Jev, a model released by TypeSafe AI in mid-September. Jev is different from the LLMs we've become used to working with. Instead of being designed to generate content, Jev is designed specifically to make decisions. You give it information, a question and a set of possible answers. Instead of generating a response, it returns a decision along with probabilities for the available choices. TypeSafe calls this new category a "System One Model" (vs LLMs which are classified as "System Two Models") and describes Jev as taking unstructured state as input and returning typed, probabilistic decisions.
For example, you might give it information about a transaction and ask, Does this transaction look suspicious? A decision model might return Suspicious: 94% / Not suspicious: 6%.
At first, I wasn't sure how interesting this really was. After all, GPT, Claude, Gemini and other LLMs can already make these kinds of decisions. But that's actually the point. Do we need to use a large, general-purpose LLM every time we need AI to make a relatively simple decision?
Why decision models could matter
Think about all the small decisions that can happen inside an AI workflow. Where should this support ticket go? Is this security event suspicious? Are these two customer records the same person? Which tool should an agent use next? Does this action require human approval? Does this AI-generated response meet our criteria?
These aren't necessarily simple enough for traditional rules. The system still needs to understand what the information means and take context into account, but it doesn't necessarily need to generate anything. That's the gap decision models are trying to fill, using AI to understand information and make a bounded decision without all of the overhead of generating a response.
This gets particularly interesting when we start talking about agents. An agent might make hundreds or thousands of small decisions as it works through a process. Today, many of those decisions may be handed to an LLM because that's the tool we have. A decision model potentially gives us another option. Let the decision model handle the smaller, repetitive decisions quickly and inexpensively, and utilize the larger LLM when you actually need its broader reasoning or generation capabilities. TypeSafe's Jev announcement and the more recent Strands Decider work both describe this type of division of labor between decision models and LLMs.
How is Jev built?
This is one of the frustrating parts of researching Jev. TypeSafe has told us quite a bit about what Jev does and says it developed a new architecture and training approach called Reinforcement Learning for Calibrated Decisions (RLCD), but it hasn't publicly disclosed enough about the underlying architecture to say exactly what Jev is.
There is another model that gives us an idea of how this type of model can be built. Laya is an open decision model created, last year, by Nandakishor M. It uses a 421M-parameter bidirectional ModernBERT-large encoder paired with a separate Transformer decision head that scores the available options rather than generating text. The Nandakishor describes it as a non-autoregressive decision model that can resolve typed decisions in a single forward pass.
Nandakishor has argued that Jev uses a similar approach. TypeSafe hasn't confirmed that, so I wouldn't assume that Jev and Laya are built the same way. But Laya does help make the idea more concrete. You can have a neural network that understands language and context without needing it to generate language.
So, what are people actually doing with Jev?
This was the next question I had when I read about Jev. It's one thing to have an interesting model. It's another to figure out where it is actually useful.
So far, people have tested Jev in several different areas. In identity resolution, security researcher Vincenzo Iozzo tested Jev against Claude models and traditional record-linkage techniques to determine whether records with slightly different names, addresses and other information actually represented the same person. In cybersecurity, Immersive Labs tested Jev against 4,134 Windows security events, using it to identify suspicious activity and group related events into investigation threads.
For AI agents, Jev has been tested as an evaluator of agent behavior, including whether an agent completed a task correctly, used its tools appropriately and produced an answer supported by the available evidence. It is also being tested as an alternative to LLM-as-a-judge, where one AI model evaluates the output or behavior of another. One recent experiment compared Jev with GPT and Claude models as judges of fixed agent runs, while LangSmith has now added Jev as an evaluator option for agent and application evaluations.
There are other examples, but these give you a good idea of the pattern. In each case, we need the AI to understand something and make a decision. We don't necessarily need it to write anything.
But, there are limitations
Jev is still AI. A structured answer doesn't magically make the answer correct. Independent testing has found differences in performance depending on the task, and even positive tests have identified situations where Jev's probabilities or decisions need to be treated carefully.
Check Point researchers have also demonstrated that a Jev-based workflow can be susceptible to prompt injection. In their experiment, attackers were able to manipulate Jev's decisions despite the structured input and output format, including when the input was explicitly identified as untrusted.
So, while a decision model may reliably give you Yes or No rather than deciding to write three paragraphs about the question, it can still confidently pick the wrong one. These are very early models, and I expect both the models and the techniques for using them to improve pretty quickly. But particularly for consequential decisions, we still need to think about confidence thresholds, validation, escalation and where a human needs to remain involved.
Wait, things just got more interesting
When Jev came out in mid-September, one of my questions was whether we were looking at the beginning of something new or just the latest AI model getting its turn in the hype cycle.
It's only been about two weeks, and now we have several more.
More importantly, they aren't just Jev clones. Some are open. Some can run locally. Some can be fine-tuned. One is intentionally small enough to sit inside an agent. Another can make decisions based on images as well as text. So I'm becoming less interested in Jev specifically and more interested in what is happening with decision models as a category.
Here are just some of them.
Kev: More control and more choices
Kev, created by Jared Palmer, is an open family of decision models ranging from 0.8B to 27B parameters. The current models are built on Qwen3.5 and Qwen3.8 and use a pointer head to score the available options without generating text. The models can be run independently, fine-tuned on your own examples and even used through the same API interface as Jev.
That means you aren't just choosing whether or not to use Kev. Instead, you choose how much model you need for the job. A small model may be enough for a simple decision. A more complicated decision may justify moving up to one of the larger models.
Kev also has an interesting approach to repeated decisions about the same information. When serving multiple questions, Kev can process the shared state once and reuse its cached representation for each question. For example, you could process a contract and then ask: What type of agreement is this? Does it contain an indemnification clause? Does it require legal review? Which department owns it? LLM platforms can also use caching techniques, so this isn't something only Kev can do, but Kev's architecture and API are specifically designed around making multiple bounded decisions against shared information.
This could make Kev particularly useful for workflows involving contracts, claims, case files or customer records where you need to make multiple decisions about the same information.
Strands Decider 2B: Small enough to live with the agent
Strands Decider 2B is deliberately small. The developers started with Qwen3.5-2B, removed the language-model head responsible for generating text and replaced it with a pointer head that scores the possible choices. The model is then fine-tuned using a LoRA adapter.
At only 2B parameters, it can run locally on relatively modest hardware. Its small size makes it well suited to operate as a component inside an agent, handling things like tool selection, routing, memory decisions, context selection, policy checks and guardrails. The Strands team specifically calls out model routing, tool selection, evaluations, guardrails, memory, context management and policy classification as areas where they are seeing early use.
Strands is also open, including the model weights, training data and training scripts. That makes it useful for understanding one way these newer decision models can actually be built.
Clef: Decision models as infrastructure
Then Cloudflare announced Clef. There are actually two models: Clef-flash, a 9B model intended for faster, latency-sensitive decisions, and Clef, a larger 27B model for more demanding decisions. Both are open models designed around the same general idea of making bounded decisions without autoregressive text generation.
One important difference with Clef is that it isn't limited to text. It can also evaluate images. Now the same idea starts expanding into things like evaluating a photo submitted with an insurance claim, classifying visual content, looking at a screenshot, or making a decision based on both an image and written information. Cloudflare also supports long context and multiple questions within a request.
Cloudflare is integrating Clef into Workers AI and its broader AI infrastructure and is also developing reinforcement-learning-based customization for decision models. This starts to show how decision models could become another layer in the AI infrastructure organizations use to build and run applications.
So how do they compare?
While similar, the models have some meaningful differences that allow them to be utilized in different ways and address more nuanced needs.
| Jev | Kev | Strands Decider 2B | Clef / Clef-flash | |
|---|---|---|---|---|
| Category | Decision model | Decision model | Decision model | Decision model |
| Primary output | Typed decisions + probabilities | Typed decisions + probabilities | Choice scores/probabilities | Typed decisions + probabilities |
| Free-form generation | No | No | No | No |
| Open weights | No | Yes | Yes | Yes |
| Size | Undisclosed | 0.8B / 4B / 9B / 27B | 2B | 27B / 9B |
| Local deployment | Not publicly offered | Yes | Yes, including modest hardware | Yes |
| Hosted option | Yes | Self-host / deploy your own endpoint | Primarily self-hosted | Yes, through Workers AI |
| Vision | No publicly documented vision support | No | No | Yes |
| Fine-tuning | Not currently user-controlled | Yes | Yes, with training recipe published | Customization / RL fine-tuning path |
| What stands out | Managed approach and more independent testing so far | Control, model-size choice and customization | Small, transparent and agent-focused | Multimodal + production infrastructure |
Sources: TypeSafe's original Jev announcement; Kev's creator-maintained repository and model documentation; the original Strands Decider announcement; and Cloudflare's original Clef announcement.
Jev has the benefit of having more independent testing behind it so far. Kev gives organizations more control over model size, hosting and customization. Strands provides a very small model that can be incorporated directly into agent architectures. Clef expands the idea beyond text while bringing Cloudflare's infrastructure into the picture.
Those are pretty different value propositions for models that, at their core, are trying to solve the same basic problem.
Where I see it getting interesting
For the last few years, we've largely talked about AI model selection as a choice between LLMs. Should I use GPT? Claude? Gemini? A smaller open model? Decision models add another dimension to that conversation. This adds even more weight and more options to deciding which parts of a workflow, or an agent's work, should be handled by which model.
- There are things where a simple rule is perfectly adequate. Use the rule.
- There are things where you need AI to understand the information and make a bounded decision, but don't need it to generate anything. Maybe that's where a decision model belongs.
- There are tasks that require broader reasoning, generation and flexibility. That's where an LLM makes sense.
- And there are particularly difficult problems where utilizing a more expensive frontier reasoning model may be justified.
- But then there are decisions where the consequences, judgment or accountability mean a human still needs to be involved
.Conceptually, you start getting something like:
Rules → Decision Models → General-Purpose LLMs → Frontier Reasoning Models → Humans
Not because every workflow has to follow that exact sequence, but because we now have more choices about how much capability we actually need for each piece of work.
Just the beginning
Decision models themselves aren't magic, and many of the ideas behind them aren't entirely new. We've had classification models, reward models, rerankers and other specialized machine-learning models for years. What I think is worth watching is how these newer decision models package semantic understanding, constrained choices, probabilities and fast inference into something developers can use much more like a general-purpose AI capability.
Two weeks isn't enough time to know exactly where this goes. But going from Jev making a splash to Kev, Strands Decider and Clef appearing in such a short period of time is enough to make anyone pay attention. Jev was announced September 15th, Kev September 17th, and Strands and Cloudflare announced their models on October 1.
And this becomes particularly important as we move further into agentic AI. An agent may make a lot of decisions while completing one task. We don't necessarily want or need our largest, most expensive model making every one of them. Decision models likely provide a new tool in our toolbox for controlling costs and creating efficiencies in agentic workflows.
Rules when rules are enough. Decision models when we need understanding and a bounded decision. LLMs when we need broader reasoning and generation. More powerful reasoning models when the complexity warrants them. Humans where judgment and accountability matter.
That makes the rise of decision models less about replacing LLMs and more about giving us another tool to use alongside them. And I suspect we're going to be hearing a lot more about them.
Key Sources
TypeSafe AI: Introducing System One Models & Jev
Laya creator announcement on Reddit
Jared Palmer: Kev repository and documentation
Strands Agents: Introducing Strands Decider 2B
Vincenzo Iozzo: Jev identity-resolution test
Immersive Labs: Testing Jev on blue-team workflows