Open decision models, measured like for like
Kahn1 4B, Clef-flash, Tev1, Laya and JEV on the same 14,663 held-out items and the 231 public JevBench items, on one GPU: the numbers, the test method, and the biases we know of, ours included.
Updated October 2, 2026
01 Β· At a glance
Clef-flash, a 9B model, is the most accurate system on our held-out set (74.8%), ahead of JEV (73.2%) and Kahn1 4B (72.4%), on the strength of Choice. On JevBench, Kahn1 4B, Jev, JevK5 and Clef-flash are within 4 points, and no gap between Kahn1 4B and any of them is significant. Tev1 (Together AI, 4B, the same base as Kahn1) is level with Kahn1 4B on the held-out set (72.0%, not a significant gap), the best on Noul (80.5%) and the best calibrated (ECE 0.060), but far behind on JevBench (76.2%, hard tier 52.3%). Laya, a 421M encoder, is far behind on both, although its training data covers the six sources of the held-out set.
| Group | Items | Kahn1 4B | Clef-flash | JEV 1.13.0 | Tev1 4B | Laya |
|---|---|---|---|---|---|---|
| All | 14,663 | 72.4% | 74.8% | 73.2% | 72.0% | 59.1% |
| Choice | 6,050 | 92.3% | 98.6% | 94.5% | 93.0% | 83.2% |
| Score | 6,210 | 51.9% | 53.7% | 52.0% | 48.4% | 29.5% |
| Noul | 2,403 | 75.5% | 69.7% | 74.4% | 80.5% | 74.8% |
ECE (0 is perfect): Kahn1 4B 0.072, Clef-flash 0.101, JEV 0.113, Tev1 0.060, Laya 0.184. Each system is measured as shipped: Kahn1 with its temperature calibration, Laya with its own, Clef-flash as a raw softmax, Tev1 through the probabilities of its option letters (which it does not present as a confidence), JEV as its API answers.
02 Β· Kahn1 held-out set, per dataset
14,663 items from the test splits of six public datasets: intents (BANKING77, MASSIVE), 5-level scales (SST-5, app reviews), entailment (RTE, SciTail). One question per request, the same options for everyone: each Choice question lists the right intent and 7 distractors seeded from the state, 8 options in all.
| Group | Items | Kahn1 4B | Clef-flash | JEV 1.13.0 | Tev1 4B | Laya |
|---|---|---|---|---|---|---|
| All | 14,663 | 72.4% | 74.8% | 73.2% | 72.0% | 59.1% |
| Choice | 6,050 | 92.3% | 98.6% | 94.5% | 93.0% | 83.2% |
| Score | 6,210 | 51.9% | 53.7% | 52.0% | 48.4% | 29.5% |
| Noul | 2,403 | 75.5% | 69.7% | 74.4% | 80.5% | 74.8% |
| BANKING77 | 3,076 | 91.8% | 99.3% | 94.8% | 94.3% | 86.1% |
| MASSIVE | 2,974 | 92.9% | 97.8% | 94.1% | 91.6% | 80.1% |
| SST-5 | 2,210 | 55.0% | 58.2% | 57.7% | 57.0% | 40.4% |
| App reviews | 4,000 | 50.2% | 51.2% | 48.9% | 43.6% | 23.5% |
| SciTail | 2,126 | 74.1% | 67.6% | 72.4% | 79.4% | 74.5% |
| RTE | 277 | 85.9% | 85.9% | 89.9% | 88.8% | 77.3% |
Paired with Kahn1 4B item by item (exact McNemar test): Clef-flash gets 1,187 items right that Kahn1 misses, Kahn1 836 that Clef-flash misses (p = 6 Γ 10β15); JEV 930 against 819 (p = 0.01); Laya 1,010 against 2,969 (p = 6 Γ 10β221); Tev1 802 against 861 (p = 0.15). Clef-flash and JEV are significantly ahead of Kahn1 4B on this held-out set.
How Noul is worded matters a lot. Asked as the bare statement (Jev's native form), the question turns into "is this true in general?": JEV drops to 48.7%, Clef-flash to 53.4%, Laya to 69.5%. The tables use The text supports this statement: {statement}, the wording that suits all three best and the one Kahn1 asks itself.
03 Β· JevBench, per tier
JevBench is an independent public benchmark for Jev-class models, with long rubrics and hard judgment calls. Clef-flash, Tev1 and Laya got the native JevBench questions and states, criteria descriptions included, as JevBench's runner sends them to Jev.
| Tier | Items | Kahn1 4B | JEV 1.13.0 | JevK5 v0.2 | Clef-flash | Tev1 4B | Laya |
|---|---|---|---|---|---|---|---|
| All | 231 | 87.4% | 86.6% | 86.1% | 83.5% | 76.2% | 57.6% |
| Standard | 72 | 97.2% | 98.6% | 95.8% | 97.2% | 97.2% | 70.8% |
| Easy | 48 | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% | 95.8% |
| Hard | 111 | 75.7% | 73.0% | 73.9% | 67.6% | 52.3% | 32.4% |
Paired with Kahn1 4B: Clef-flash 12 against 21 (p = 0.16, not significant), Laya 8 against 77 (p = 3 Γ 10β15), Tev1 7 against 33 (p = 4 Γ 10β5), Jev 11 against 13 (p = 0.84), JevK5 10 against 13 (p = 0.68). Three runs on the same items: ours for Kahn1, Clef-flash, Tev1 and Laya, JevBench's for Jev, JevK5's authors' for JevK5 (JevBench's own run gives JevK5 v0.2 85.3%).
04 Β· How each system was run
- Same items, same options, same labels. Choice over the same 8 options as Kahn1, Score over the same levels, Noul in the wording above. Questions are sent in Jev's fields (
type,instructions,criteria), Clef-flash's and Laya's native format. - Kahn1 4B:
Okura66/Kahn1-Qwen3.5-4B, vLLM, bf16, k = 3 option orders averaged, temperature calibration, its saved predictions. - Clef-flash:
Cloudflare/clef-flashthrough its release code (joint_schema_model.py), transformers, a softmax per question. Weights in int8 (bitsandbytes): in bf16 its 18.8 GB do not fit in our card's 16 GB. Text only. - Tev1:
togethercomputer/Tev1-4B-experimental, vLLM, bf16, its recommended system prompt and JSON decision (state, question, letter-labelled options), thinking off. Its answer is the most likely letter at the first token, the one it generates greedily; the letters' probabilities serve as its confidence. Score goes through its options, in level order; Noul through two options, yes and no. - Laya:
convaiinnovations/laya0.3.23, its recommendedRouter(which picks the English or the multilingual checkpoint), its shipped calibration. - JEV 1.13.0: TypeSafe's API, called by us with the same items (held-out set); on JevBench, the outcomes JevBench publishes.
- Hardware: one RTX 5070 Ti (16 GB), WSL2. Statistics: paired exact McNemar test, ECE over 15 bins of pmax. One run per system: all are deterministic at inference.
05 Β· Known biases and limits
- The held-out set is ours. We chose its six sources and its shape (8 options on Choice). They were kept out of Kahn1's training, but Kahn1 was trained on neighbouring tasks: intents (CLINC150, which has banking intents), sentiment (Yelp, Amazon, IMDB), NLI (MNLI, ANLI, WANLIβ¦).
- Laya was trained on the six sources of the held-out set (BANKING77, MASSIVE, SST-5, SciTail, RTE, app_reviews, per its model card): its held-out figures are in-distribution, not zero-shot, and still behind.
- Clef-flash's training data is not published ("internal synthetic datasets"). BANKING77 is among its provider's benchmarks. See the next section.
- Tev1 was trained on BANKING77 and SST-5 (and MultiNLI, BoolQ, AG News, plus synthetic rules, per its
DATA_SOURCES.md): its figures on those two sources are in-distribution. Not on JevBench, which it says it did not use. Its weights licence is still "being finalized". - Clef-flash ran in int8, not in bf16 as released: its figures may move a little in bf16.
- Laya reads 512 tokens with its English checkpoint: 57 of the 231 JevBench states are truncated. Read in full by its multilingual checkpoint (8,192 tokens), it scores lower (46.8%): truncation does not explain its score.
- Shipped settings: Kahn1 averages 3 option orders and calibrates, the others answer in one pass, Tev1 with the options in the order given. ECE therefore compares systems as shipped, not after a common recalibration.
- Three runners on JevBench; only JevBench's public half could be tested here.
06 Β· Did a system see the test sources?
Two tests. First a probe: 2,000 items drawn from the train splits of five sources, asked exactly like the held-out set. A model that memorised those items scores higher on train than on test; a model that never saw the source scores the same on both. Kahn1, which saw none of these sources, is the control; Noul is asked as the bare statement for Clef-flash and Laya, on both splits.
| Source | Kahn1 4B, test β train | Clef-flash, test β train | Tev1 4B, test β train | Laya, test β train |
|---|---|---|---|---|
| BANKING77 | 91.8% β 91.5% (-0.3 pt) | 99.3% β 99.2% (-0.1 pt) | 94.3% β 94.8% (+0.5 pt) | 86.1% β 84.7% (-1.4 pt) |
| MASSIVE | 92.9% β 93.4% (+0.5 pt) | 97.8% β 98.7% (+0.9 pt) | 91.6% β 92.3% (+0.7 pt) | 80.1% β 81.3% (+1.2 pt) |
| SST-5 | 55.0% β 55.4% (+0.4 pt) | 58.2% β 60.5% (+2.3 pt) | 57.0% β 57.9% (+0.9 pt) | 40.4% β 36.6% (-3.8 pt) |
| SciTail | 74.1% β 72.0% (-2.2 pt) | 50.3% β 48.2% (-2.1 pt) | 60.3% β 58.1% (-2.3 pt) | 68.3% β 66.8% (-1.6 pt) |
| RTE | 85.9% β 87.3% (+1.4 pt) | 77.3% β 77.1% (-0.2 pt) | 81.6% β 81.9% (+0.3 pt) | 78.0% β 73.7% (-4.3 pt) |
No system shows a gap beyond the control's, neither Laya nor Tev1, which say they were trained on these sources (all of them for Laya, BANKING77 and SST-5 for Tev1). The probe detects memorisation, not exposure: a model trained on a train split without overfitting does as well on the test split. It does not prove that a system never saw this data.
Then, the gain over the base model, at k = 1 without calibration: what fine-tuning adds to the raw model.
| Group | Qwen3.5-4B (base) | Kahn1 4B, k = 1 | Tev1 4B | Qwen3.5-9B (base, fp8) | Clef-flash |
|---|---|---|---|---|---|
| All | 62.2% | 72.0% | 72.0% | 66.3% | 74.8% |
| Choice | 88.3% | 92.1% | 93.0% | 92.0% | 98.6% |
| Score | 39.3% | 51.2% | 48.4% | 45.2% | 53.7% |
| Noul | 55.7% | 75.2% | 80.5% | 55.9% | 69.7% |
| BANKING77 | 87.9% | 91.2% | 94.3% | 91.8% | 99.3% |
| MASSIVE | 88.7% | 93.1% | 91.6% | 92.3% | 97.8% |
On BANKING77, Clef-flash removes 92% of its base model's errors (from 8.2% to 0.7% errors), and 71% on MASSIVE; Kahn1 4B removes 27% and 39% (Kahn1 saw neither), and Tev1, which saw BANKING77 but not MASSIVE, 53% and 26%. A gain that large on intents is what training on intent data gives, whether these datasets or close ones, and these two tests cannot tell which. The honest reading: Clef-flash's lead on Choice is real on these items, but we cannot claim it is zero-shot. The base models are read with Kahn1's prompt format, the 9B in fp8.
07 Β· Data and reproduction
Per-item outcomes: open_models_heldout.csv (held-out set, right or wrong and confidence for each system) and open_models_jevbench.csv. The item index points into data/eval.jsonl, which training/build_dataset.py rebuilds identically.
Scripts: scripts/jev_models_eval.py (Clef-flash, Laya, report), scripts/contamination_probe.py (probe), scripts/eval_base.py (base models). Reports: OPEN_DECISION_MODELS.md, CONTAMINATION_PROBE.md.
Head to head: Kahn1 vs Clef-flash, Kahn1 vs Laya. Think a system was run wrong? Open an issue and we will re-run it.