System One models
A System One model answers typed questions about a text in one forward pass, and returns each answer with a probability for every allowed option. TypeSafe's Jev made the term popular in September 2026; open models now do the same.
Updated October 2, 2026
01 · Definition
The name comes from Daniel Kahneman's System 1 (Thinking, Fast and Slow, 2011): fast, cheap judgments that settle most cases and leave the hard ones to a slower System 2 (a frontier model or a person). TypeSafe writes “System One”; you will also read “System 1”, “decision model” and “Jev-class model”.
Such a model takes a state (a message, a document, a ticket) and typed questions. For each it returns a typed value and a probability distribution: nothing is generated, so nothing has to be parsed, and an answer outside the schema cannot happen.
| Question type | What you ask | What comes back |
|---|---|---|
| Choice | A fixed list of options | One option + a probability for each |
| Score | Ordered levels | A level + an expected value over the levels |
| Noul (yes / no) | A statement to check | The probability that the text supports it |
02 · How it works
The prompt holds the state, the question and the options, each tagged with a letter. The model runs one forward pass; at the answer position, only the logits of the option tokens (A, B, C… or yes / no) are kept, and the softmax turns them into probabilities. For a Score, the expected value Σ i · pi gives a continuous value.
Two settings make those probabilities usable. Permutation debiasing asks the question with the options in k orders and averages the results. Calibration (one temperature per question type) aligns confidence with accuracy, measured by the ECE. On a long state, prefix caching also reuses the state's computation from one question to the next; on a short state the engine recomputes it for each question, and latency grows with the number of questions.
| System One model | LLM generating JSON | |
|---|---|---|
| Output | A typed value with its distribution | Text to validate and parse |
| Format errors | Impossible by construction | Rare with structured outputs, possible otherwise |
| Probabilities | One per option, always | Logprobs depending on the provider, often none |
| What drives latency | One pass per question, over the state | The length of the generated answer |
| What drives cost | Input tokens | Input and output tokens |
03 · The families
Seven families turn a text into a typed decision; only decision models combine a guaranteed format, a probability per option and a single pass. Product by product, with licences and sources, see the landscape.
| Family | Examples | Licence | Where it runs |
|---|---|---|---|
| Hosted decision models | Jev (TypeSafe), d1 (Liquid AI), Decisions API (OpenAI, limited preview) | Closed | API |
| Open decision models | Clef (Cloudflare), JevK5, Kahn1, Tev1 (Together), Laya | Mostly Apache 2.0 | Local, sometimes also hosted |
| LLMs with structured outputs | OpenAI, Anthropic, Google | Closed | API |
| Constrained decoding | Outlines, XGrammar, vLLM, llama.cpp, Guidance | Apache 2.0 or MIT | Local, with any model |
| Cloud classification | Google Natural Language, AWS Comprehend, Azure Language | Closed | API |
| Zero-shot encoders | bart-large-mnli, DeBERTa-v3 zeroshot, GLiClass, SetFit | MIT or Apache 2.0 | CPU |
| Guard and judge models | Llama Guard 4, ShieldGemma, Granite Guardian, Qwen3Guard, Prometheus 2 | Mixed | Local |
04 · How to choose
- Can your data leave your infrastructure? No: an open model, run locally. Yes: a hosted API becomes an option.
- Which language? Jev is trained first for English, per its documentation; check the language of each model.
- Which licence? Apache 2.0 and MIT allow commercial use; some licences (Qwen Research, CC BY-NC) limit it.
- Does calibration hold on your data? Measure the ECE on a few hundred labelled examples before you set a threshold.
- How many options, how long a text? The number of options and the context window differ from one model to the next.
- What throughput, what hardware? A GPU for production; a CPU is enough to try it, at seconds per question.
- Is an encoder enough? For short, fixed intents, a zero-shot encoder on a CPU is often faster and good enough.
05 · How they are measured
JevBench is the independent benchmark of the category: states and bounded rubrics, a typed answer, four axes (accuracy, calibration, speed, cost). Vendors also cite their own indices: Cloudflare and Liquid AI each report beating Jev on a “Decision Index”, figures not reproduced here.
A comparison only means something like for like: the same items, the same options, the same labels. That is the rule of our benchmarks: on 14,663 held-out items, Jev 1.13.0 scores 73.2% and Kahn1 4B 72.4%. On the 231 public JevBench items, Jev scores 86.6% in JevBench's run, Kahn1 4B 87.4% in ours (k = 3, calibrated; not a significant gap, p = 0.84) and JevK5 v0.2 85.3% in JevBench's run (86.1% in its authors' own run): three separate runs on the same items.
06 · Limits
- They are not reasoning models: computing a date, a duration or an amount in one pass stays a weak spot.
- Calibration holds on the distribution it was fitted on: recalibrate on your domain.
- Confidence is not the probability of being right: set thresholds on your own data.
- Regulation still applies (EU AI Act, sector rules): keep a person or a larger model on the low-confidence path.
For Kahn1 in detail, see the caveats.
07 · Key terms
- System One model. A model that answers typed questions about a text in one forward pass, with a probability for each allowed answer. Also called a decision model or a Jev-class model.
- Choice. A question with a fixed list of options; the answer is one option and a probability for each.
- Score. A question on ordered levels (low, medium, high); the answer is a level and an expected value over the levels.
- Noul. Jev's name for a yes / no question: does the text support this statement? Kahn1 uses the same name.
- Logit. The raw score a language model gives each token before the softmax turns scores into probabilities.
- Calibration. How well probabilities match reality: of the answers given at 80% confidence, about 80% should be right.
- ECE. Expected calibration error: the average gap between confidence and accuracy, over bins of confidence. 0 is perfect.
- Permutation debiasing. Asking the same question with the options in k different orders and averaging the probabilities, so the position of an option does not sway the answer.
- Prefix caching. Reusing the computation of a shared prompt prefix (the state) across the questions asked about it, when the engine can: vLLM caches Kahn1 4B's prefix in 528-token blocks, so only states longer than a block benefit.