KAHN1
EN FR

Open decision models, measured like for like

Kahn1 4B, Clef-flash, Tev1, Laya and JEV on the same 14,663 held-out items and the 231 public JevBench items, on one GPU: the numbers, the test method, and the biases we know of, ours included.

Updated October 2, 2026

01 Β· At a glance

Clef-flash, a 9B model, is the most accurate system on our held-out set (74.8%), ahead of JEV (73.2%) and Kahn1 4B (72.4%), on the strength of Choice. On JevBench, Kahn1 4B, Jev, JevK5 and Clef-flash are within 4 points, and no gap between Kahn1 4B and any of them is significant. Tev1 (Together AI, 4B, the same base as Kahn1) is level with Kahn1 4B on the held-out set (72.0%, not a significant gap), the best on Noul (80.5%) and the best calibrated (ECE 0.060), but far behind on JevBench (76.2%, hard tier 52.3%). Laya, a 421M encoder, is far behind on both, although its training data covers the six sources of the held-out set.

GroupItemsKahn1 4BClef-flashJEV 1.13.0Tev1 4BLaya
All14,66372.4%74.8%73.2%72.0%59.1%
Choice6,05092.3%98.6%94.5%93.0%83.2%
Score6,21051.9%53.7%52.0%48.4%29.5%
Noul2,40375.5%69.7%74.4%80.5%74.8%

ECE (0 is perfect): Kahn1 4B 0.072, Clef-flash 0.101, JEV 0.113, Tev1 0.060, Laya 0.184. Each system is measured as shipped: Kahn1 with its temperature calibration, Laya with its own, Clef-flash as a raw softmax, Tev1 through the probabilities of its option letters (which it does not present as a confidence), JEV as its API answers.

02 Β· Kahn1 held-out set, per dataset

14,663 items from the test splits of six public datasets: intents (BANKING77, MASSIVE), 5-level scales (SST-5, app reviews), entailment (RTE, SciTail). One question per request, the same options for everyone: each Choice question lists the right intent and 7 distractors seeded from the state, 8 options in all.

GroupItemsKahn1 4BClef-flashJEV 1.13.0Tev1 4BLaya
All14,66372.4%74.8%73.2%72.0%59.1%
Choice6,05092.3%98.6%94.5%93.0%83.2%
Score6,21051.9%53.7%52.0%48.4%29.5%
Noul2,40375.5%69.7%74.4%80.5%74.8%
BANKING773,07691.8%99.3%94.8%94.3%86.1%
MASSIVE2,97492.9%97.8%94.1%91.6%80.1%
SST-52,21055.0%58.2%57.7%57.0%40.4%
App reviews4,00050.2%51.2%48.9%43.6%23.5%
SciTail2,12674.1%67.6%72.4%79.4%74.5%
RTE27785.9%85.9%89.9%88.8%77.3%

Paired with Kahn1 4B item by item (exact McNemar test): Clef-flash gets 1,187 items right that Kahn1 misses, Kahn1 836 that Clef-flash misses (p = 6 Γ— 10βˆ’15); JEV 930 against 819 (p = 0.01); Laya 1,010 against 2,969 (p = 6 Γ— 10βˆ’221); Tev1 802 against 861 (p = 0.15). Clef-flash and JEV are significantly ahead of Kahn1 4B on this held-out set.

How Noul is worded matters a lot. Asked as the bare statement (Jev's native form), the question turns into "is this true in general?": JEV drops to 48.7%, Clef-flash to 53.4%, Laya to 69.5%. The tables use The text supports this statement: {statement}, the wording that suits all three best and the one Kahn1 asks itself.

03 Β· JevBench, per tier

JevBench is an independent public benchmark for Jev-class models, with long rubrics and hard judgment calls. Clef-flash, Tev1 and Laya got the native JevBench questions and states, criteria descriptions included, as JevBench's runner sends them to Jev.

TierItemsKahn1 4BJEV 1.13.0JevK5 v0.2Clef-flashTev1 4BLaya
All23187.4%86.6%86.1%83.5%76.2%57.6%
Standard7297.2%98.6%95.8%97.2%97.2%70.8%
Easy48100.0%100.0%100.0%100.0%100.0%95.8%
Hard11175.7%73.0%73.9%67.6%52.3%32.4%

Paired with Kahn1 4B: Clef-flash 12 against 21 (p = 0.16, not significant), Laya 8 against 77 (p = 3 Γ— 10βˆ’15), Tev1 7 against 33 (p = 4 Γ— 10βˆ’5), Jev 11 against 13 (p = 0.84), JevK5 10 against 13 (p = 0.68). Three runs on the same items: ours for Kahn1, Clef-flash, Tev1 and Laya, JevBench's for Jev, JevK5's authors' for JevK5 (JevBench's own run gives JevK5 v0.2 85.3%).

04 Β· How each system was run

  • Same items, same options, same labels. Choice over the same 8 options as Kahn1, Score over the same levels, Noul in the wording above. Questions are sent in Jev's fields (type, instructions, criteria), Clef-flash's and Laya's native format.
  • Kahn1 4B: Okura66/Kahn1-Qwen3.5-4B, vLLM, bf16, k = 3 option orders averaged, temperature calibration, its saved predictions.
  • Clef-flash: Cloudflare/clef-flash through its release code (joint_schema_model.py), transformers, a softmax per question. Weights in int8 (bitsandbytes): in bf16 its 18.8 GB do not fit in our card's 16 GB. Text only.
  • Tev1: togethercomputer/Tev1-4B-experimental, vLLM, bf16, its recommended system prompt and JSON decision (state, question, letter-labelled options), thinking off. Its answer is the most likely letter at the first token, the one it generates greedily; the letters' probabilities serve as its confidence. Score goes through its options, in level order; Noul through two options, yes and no.
  • Laya: convaiinnovations/laya 0.3.23, its recommended Router (which picks the English or the multilingual checkpoint), its shipped calibration.
  • JEV 1.13.0: TypeSafe's API, called by us with the same items (held-out set); on JevBench, the outcomes JevBench publishes.
  • Hardware: one RTX 5070 Ti (16 GB), WSL2. Statistics: paired exact McNemar test, ECE over 15 bins of pmax. One run per system: all are deterministic at inference.

05 Β· Known biases and limits

  • The held-out set is ours. We chose its six sources and its shape (8 options on Choice). They were kept out of Kahn1's training, but Kahn1 was trained on neighbouring tasks: intents (CLINC150, which has banking intents), sentiment (Yelp, Amazon, IMDB), NLI (MNLI, ANLI, WANLI…).
  • Laya was trained on the six sources of the held-out set (BANKING77, MASSIVE, SST-5, SciTail, RTE, app_reviews, per its model card): its held-out figures are in-distribution, not zero-shot, and still behind.
  • Clef-flash's training data is not published ("internal synthetic datasets"). BANKING77 is among its provider's benchmarks. See the next section.
  • Tev1 was trained on BANKING77 and SST-5 (and MultiNLI, BoolQ, AG News, plus synthetic rules, per its DATA_SOURCES.md): its figures on those two sources are in-distribution. Not on JevBench, which it says it did not use. Its weights licence is still "being finalized".
  • Clef-flash ran in int8, not in bf16 as released: its figures may move a little in bf16.
  • Laya reads 512 tokens with its English checkpoint: 57 of the 231 JevBench states are truncated. Read in full by its multilingual checkpoint (8,192 tokens), it scores lower (46.8%): truncation does not explain its score.
  • Shipped settings: Kahn1 averages 3 option orders and calibrates, the others answer in one pass, Tev1 with the options in the order given. ECE therefore compares systems as shipped, not after a common recalibration.
  • Three runners on JevBench; only JevBench's public half could be tested here.

06 Β· Did a system see the test sources?

Two tests. First a probe: 2,000 items drawn from the train splits of five sources, asked exactly like the held-out set. A model that memorised those items scores higher on train than on test; a model that never saw the source scores the same on both. Kahn1, which saw none of these sources, is the control; Noul is asked as the bare statement for Clef-flash and Laya, on both splits.

SourceKahn1 4B, test β†’ trainClef-flash, test β†’ trainTev1 4B, test β†’ trainLaya, test β†’ train
BANKING7791.8% β†’ 91.5% (-0.3 pt)99.3% β†’ 99.2% (-0.1 pt)94.3% β†’ 94.8% (+0.5 pt)86.1% β†’ 84.7% (-1.4 pt)
MASSIVE92.9% β†’ 93.4% (+0.5 pt)97.8% β†’ 98.7% (+0.9 pt)91.6% β†’ 92.3% (+0.7 pt)80.1% β†’ 81.3% (+1.2 pt)
SST-555.0% β†’ 55.4% (+0.4 pt)58.2% β†’ 60.5% (+2.3 pt)57.0% β†’ 57.9% (+0.9 pt)40.4% β†’ 36.6% (-3.8 pt)
SciTail74.1% β†’ 72.0% (-2.2 pt)50.3% β†’ 48.2% (-2.1 pt)60.3% β†’ 58.1% (-2.3 pt)68.3% β†’ 66.8% (-1.6 pt)
RTE85.9% β†’ 87.3% (+1.4 pt)77.3% β†’ 77.1% (-0.2 pt)81.6% β†’ 81.9% (+0.3 pt)78.0% β†’ 73.7% (-4.3 pt)

No system shows a gap beyond the control's, neither Laya nor Tev1, which say they were trained on these sources (all of them for Laya, BANKING77 and SST-5 for Tev1). The probe detects memorisation, not exposure: a model trained on a train split without overfitting does as well on the test split. It does not prove that a system never saw this data.

Then, the gain over the base model, at k = 1 without calibration: what fine-tuning adds to the raw model.

GroupQwen3.5-4B (base)Kahn1 4B, k = 1Tev1 4BQwen3.5-9B (base, fp8)Clef-flash
All62.2%72.0%72.0%66.3%74.8%
Choice88.3%92.1%93.0%92.0%98.6%
Score39.3%51.2%48.4%45.2%53.7%
Noul55.7%75.2%80.5%55.9%69.7%
BANKING7787.9%91.2%94.3%91.8%99.3%
MASSIVE88.7%93.1%91.6%92.3%97.8%

On BANKING77, Clef-flash removes 92% of its base model's errors (from 8.2% to 0.7% errors), and 71% on MASSIVE; Kahn1 4B removes 27% and 39% (Kahn1 saw neither), and Tev1, which saw BANKING77 but not MASSIVE, 53% and 26%. A gain that large on intents is what training on intent data gives, whether these datasets or close ones, and these two tests cannot tell which. The honest reading: Clef-flash's lead on Choice is real on these items, but we cannot claim it is zero-shot. The base models are read with Kahn1's prompt format, the 9B in fp8.

07 Β· Data and reproduction

Per-item outcomes: open_models_heldout.csv (held-out set, right or wrong and confidence for each system) and open_models_jevbench.csv. The item index points into data/eval.jsonl, which training/build_dataset.py rebuilds identically.

Scripts: scripts/jev_models_eval.py (Clef-flash, Laya, report), scripts/contamination_probe.py (probe), scripts/eval_base.py (base models). Reports: OPEN_DECISION_MODELS.md, CONTAMINATION_PROBE.md.

Head to head: Kahn1 vs Clef-flash, Kahn1 vs Laya. Think a system was run wrong? Open an issue and we will re-run it.