AI News

Decision AI Models: Jev, Rivals, and Real Results Without Extra Tokens

Decision AI models return typed decisions with probabilities, not text. We unpack Jev, benchmarks, pricing, use cases, competitors, and how to choose.

2026-10-03 ·Hai Anton

Decision AI models produce choices, not text. You send a state with typed questions, and the model returns options, scores, or probabilities your code can branch on. The category went mainstream after TypeSafe AI launched Jev, built over two stealth years and branded a “System One” model after Kahneman’s fast System 1. Within three weeks, rivals and open‑source reproductions appeared. Here is how they work, what the numbers show, pricing, use cases, and how to pick a model.

What are Decision AI models, and why now?

They return a decision, not a paragraph. You ask typed questions and get a Choice, a Score, or a Yes/No probability your code can act on.

The category broke out after Jev from TypeSafe AI. The team built it in stealth for two years and calls it a “System One” model, nodding to Kahneman’s fast, intuitive System 1 thinking. The launch triggered a wave.

Within three weeks, Fastino Labs shipped two competing models. Open‑source developers released several Jev‑style reproductions in parallel. A distinct class with its own ecosystem formed quickly.

The rule is simple: if your code needs a bounded answer for branching, decision models fit. If a human must read the output, use an LLM. Next, we cover Jev’s mechanics, benchmarks, pricing, and practical cases.

Which matters most for you — speed, cost, or ecosystem? That question will guide your first pilot.

How does Jev work, and what does it cost?

Jev takes a “state” and one or more questions. It supports three primitives — Choice, Score, and Noul — and evaluates all questions in parallel and in isolation.

Per TypeSafe docs, Choice selects a single option from a list with probabilities and confidence. Jev supports up to 255 options. Score rates the state on an ordered rubric, also with probabilities and confidence. Noul is a 0 to 1 probability that a statement is true.

Each question is evaluated independently against the same state. Adding more questions barely changes latency. Jev never generates strings, so, per TypeSafe, a type mismatch cannot occur.

Under the hood, TypeSafe describes a new architecture, a parallel sampler, and a training method called Reinforcement Learning for Calibrated Decisions (RLCD). RLHF optimizes for human preference, while RLCD targets calibrated probabilities where higher confidence aligns with higher accuracy.

Pricing is the headline. Jev costs $0.042 per million input tokens, with free output. OpenRouter lists a 32K context window. TypeSafe reports end‑to‑end latency from 70 to 500 ms.

What do TypeSafe’s benchmarks show?

In workflow evals, Jev matched Sonnet 5 on accuracy at a fraction of cost and latency. It still trails the top frontier configuration by 6.3 points.

TypeSafe built evaluations across four tasks: security incidents, agent trace observability, invoice processing, and customer service. Reference labels came from averaging GPT‑6 Astra and Claude Fable 5.1 at high thinking.

The results were: Jev at 67.8% mean accuracy, $0.0004 per case, and 0.4 seconds. In the same workflow, Claude Sonnet 5 posted 67.8%, $0.1174 per case, and 78.1 seconds. The best comparison model (OpenAI “sol”) showed 74.1%, $0.0836, and 23.3 seconds.

Per task, the picture mixed. Jev scored 76.0% on customer service but only 61.8% on invoice processing. The takeaway: cost and latency shine; accuracy depends on context.

These figures come from TypeSafe’s workflow evaluations and reflect that methodology.

Where do decision models fit, and where don’t they?

The rule holds: use them when code needs a bounded answer for branching. Use an LLM when a person must read the output.

Agent control flow is a natural fit. Choose the next tool or subagent, or decide “continue, retry, ask the user, or stop” with a single Choice. Route requests by difficulty: easy to a cheap model, hard to a frontier one. Fastino lists routing by destination, complexity, or escalation level.

Classification and triage are another big area. Ticket routing, email triage, and intent detection make up much of Fastino’s 17‑dataset benchmark. Simon Willison highlights spam detection, label suggestions, and prioritization as natural fits. In security, TypeSafe’s simplest workflow decides to close an alert, pass it to an analyst, or contain it.

Verification and safety also benefit. Use Jev to score prompts, traces, and outputs for guardrails and jailbreak detection. GLiNER2.5‑Decide decodes “safety” and “harm type” jointly so answers do not contradict. OpenRouter outlines verified cascades: draft cheaply, check with Jev, and escalate only on failures.

Evaluation and observability follow suit. Arize and Langfuse run Jev‑as‑a‑judge evaluators on traces. In CI/CD, Buddy lets pipelines score, classify, or gate runs with a Jev action. In search and data work, you can rerank 100 BM25 candidates in one parallel call, map‑reduce large datasets into features, and prune context before an LLM.

Real‑time use cases are rising too. TypeSafe demoed Jev playing Doom and Wikiracing. Sub‑second decisions keep UX responsive. The Doom demo ran at about 10 queries per second, with an estimated cost near $7 per hour.

Where not to use them? When you need generated text, summaries, or explanations. When exact arithmetic, counting, or date math is required. TypeSafe’s Jev 1.13 jaggedness guide flags all three. And when decisions affect livelihoods, such as hiring. Willison warns that hidden bias is hard to inspect.

Who competes, and how should you choose today?

Several lines are already prominent. Jev 1.13 by TypeSafe is a closed, hosted API at $0.042 per million input tokens, free output, with a new architecture, a parallel sampler, and RLCD, up to 255 options in Choice, and 70–500 ms latency. GLiDE from Fastino is a closed API with a 40K context and “adaptive thinking” on uncertain cases, returning a selected action, confidence, and option probabilities.

GLiNER2.5‑Decide ships open weights under Apache 2.0. It is a 340M DeBERTa‑v3‑large encoder for typed decisions with joint constraints, plus spans and relations. Reported p50 is 38.3 ms on a V100 and 167.3 ms on a 48‑vCPU CPU, and it runs locally. Laya by Convai is also Apache 2.0, at 421M ModernBERT‑large (English) or 322M mmBERT‑base (multilingual), supports Choice, Score, and Noul across 100+ languages, about 33 ms per question on GPU, and runs locally.

Reproductions on Qwen exist as well. JevK5 is a 4B model on Qwen3.5‑4B with merged LoRA, returning a probability for every option in one pass and running on a GPU locally. OpenJev (MIT) reads option logits on a frozen Qwen3.5‑4B and is 5.21x faster than autoregressive JSON. kev‑0.5b adds a LoRA and readout head to Qwen2.5‑0.5B with a Jev‑compatible API, supports Noul and Choice 2–255 and Score, and clocks around ~160 ms for six questions on an Apple M5, running on a laptop.

Head‑to‑head results come from different test suites and should not be compared across rows. GLiDE reports 64.81 on Decision Index 0.2.1. On CLadder, GLiDE hits 88.7% and Jev 72.6%. On CRUXEval, GLiDE is 92.6% and Jev 73.0%. On Fast Decisions (17 datasets), GLiNER2.5‑Decide leads at 60.1%, with JevK5 at 57.5%, SemIf at 56.4%, GLiFormer at 49.0%, and Laya at 46.6%. On a 102‑row TypeSafe eval subset, Jev reports 0.883 and OpenJev 0.845 balanced accuracy.

Two caveats matter. Fastino chose both the tests and the opponents for its GLiDE and GLiNER2.5‑Decide claims. Its Fast Decisions comparison uses JevK5, an open reproduction, not TypeSafe’s Jev. Always verify on your data.

How to choose? Pick Jev for a managed API with the broadest ecosystem support today. Choose GLiDE to probe harder, reasoning‑heavy decisions. Choose GLiNER2.5‑Decide for air‑gapped deployment, fine‑tuning, or answers that must obey cross‑question rules. Pick Laya for multilingual needs or very low per‑question latency. And use kev, OpenJev, or JevK5 to experiment locally or avoid vendor lock‑in.

Why the trend now? The idea is not new: classifiers and rerankers have made decisions for years. Decision Transformer (2021) framed RL as sequence modeling. DeepMind’s Gato (2022) acted across many tasks. A 2023 survey even called “large decision models” the next step. In 2026, packaging changed the game: agents need cheap judgments, calibration unlocks automation, the ecosystem moved quickly (Vercel, OpenRouter, Arize, Langfuse, Buddy), and competition landed immediately — Fastino shipped GLiNER2.5‑Decide on September 24 and GLiDE on September 30, while kev, OpenJev, and JevK5 appeared in parallel.

Based on the provided brief.

Ready to automate your store?

We'll analyze your workflows, find the bottlenecks, and propose a concrete automation plan. First consultation is free.

Message us on Telegram →
Hai Anton
Hai Anton

Founder of HAIQ — AI Automation Agency. Founder of HAIQ. I build automations and AI solutions for Ukrainian e-commerce on n8n. I write about automation, chatbots, and AI for business.