KAHN1
EN FR

Benchmarks

Every number behind the home page: Kahn1 4B and Kahn1 3B against JEV on the same items, Kahn1's own held-out results, latency, and how to reproduce them.

01 · At a glance

Two benchmarks, both paired item by item. On the Kahn1 held-out set JEV was run by us through its API, with the same states, prompts and options as Kahn1. On JevBench, Jev's results are the ones JevBench itself publishes. The chart shows Kahn1 4B and Kahn1 3B; the tables below add the per-source figures, and JevK5 on JevBench.

Other open models. Run by us on the same items and the same GPU, Cloudflare's Clef-flash (9B, in int8) is the most accurate on the held-out set, 74.8% against 72.4% for Kahn1 4B and 73.2% for JEV, and level with Kahn1 4B on JevBench (83.5%, not a significant gap); Laya (421M) scores 59.1% and 57.6%. Every table, the test method and the known biases: open decision models, measured like for like.

Correction. An earlier version of this page asked Choice over 8 options for Kahn1 but over every intent for JEV, which made Kahn1 look ahead on Choice and overall. That comparison was unequal and has been corrected: over the same options, JEV is ahead on Choice and overall. Both conditions, side by side (3B).

Kahn1 4BKahn1 3BJEV 1.13.0 · TypeSafe API
KAHN1 HELD-OUT
14,663 items · same options · JEV run by us through its API
0255075100
929194
525152
756774
727073
ChoiceScoreNoulAll
JEVBENCH PUBLIC
231 items · Jev scored by JevBench
0255075100
100100100
978599
764273
876887
EasyStandardHardAll
TABLE VIEW
BenchmarkGroupItemsKahn1 4BKahn1 3BJEV
KAHN1 HELD-OUTChoice6,05092.3%90.9%94.5%
KAHN1 HELD-OUTScore6,21051.9%51.4%52.0%
KAHN1 HELD-OUTNoul2,40375.5%67.0%74.4%
KAHN1 HELD-OUTAll14,66372.4%70.3%73.2%
JEVBENCH PUBLICEasy48100.0%100.0%100.0%
JEVBENCH PUBLICStandard7297.2%84.7%98.6%
JEVBENCH PUBLICHard11175.7%42.3%73.0%
JEVBENCH PUBLICAll23187.4%67.5%86.6%

02 · Kahn1 vs JEV on the Kahn1 held-out set

14,663 items from six public datasets, none seen in training, in the task families Kahn1 was trained on: intents (banking77, MASSIVE), 5-level scales (SST-5, app reviews) and entailment (RTE, SciTail). One question per request, and the same options for every system: each Choice question lists the right intent and 7 distractors seeded from the state, 8 options in all. Kahn1 4B and Kahn1 3B: k = 3, temperature calibration. JEV: jev-1.13.0, answered all 14,663 without error. ECE: 15 bins on pmax, 0 is perfect.

DatasetItemsKahn1 4BKahn1 3BJEV4B ECE3B ECEJEV ECE
banking773,07691.8%91.1%94.8%0.0180.0260.024
MASSIVE2,97492.9%90.8%94.1%0.0130.0110.017
RTE27785.9%82.7%89.9%0.0480.1050.034
SciTail2,12674.1%65.0%72.4%0.1410.2480.077
SST-52,21055.0%52.4%57.7%0.0520.0980.178
App reviews4,00050.2%50.9%48.9%0.1460.0960.252
Choice6,05092.3%90.9%94.5%0.0150.0170.019
Score6,21051.9%51.4%52.0%0.1120.0970.226
Noul2,40375.5%67.0%74.4%0.1270.2320.071
All14,66372.4%70.3%73.2%0.0720.0820.113

JEV is ahead on Choice and overall and level on Score (52.0% against 51.9%); Kahn1 4B is ahead on Noul (75.5% against 74.4%, a gap not tested for significance). The 4B is ahead of the 3B on every primitive and the better calibrated overall.

Choice over every intent

The table above asks Choice over 8 options. Asked over every intent instead (77 for banking77, 60 for MASSIVE), on the same 1,184 items, every system drops, Kahn1 more than JEV. Kahn1 goes through its two-stage router there: one Noul per option keeps the 10 best, then a Choice picks among them (k = 1).

Options per questionItemsKahn1 4BKahn1 3BJEV
8 (the right one + 7 seeded distractors)6,05092.3%90.9%94.5%
every intent (77 / 60)1,18471.6%68.4%79.1%

How the Noul question is worded

JEV's native Noul form is the bare statement as instructions, which is also what JevBench sends. JEV then judges whether the statement is true in general, not whether the text supports it: on SciTail, whose negatives are statements the text does not support, it answers yes to 93% of items. Kahn1's prompt asks whether the statement is "true for this state", so the Noul items were re-run with The text supports this statement: {statement}. The tables above use that wording, the fairer one for JEV.

DatasetItemsKahn1 4BKahn1 3BJEV, bare statementJEV, “supported” wording
RTE27785.9%82.7%82.3%89.9%
SciTail2,12674.1%65.0%44.4%72.4%
Noul2,40375.5%67.0%48.7%74.4%

03 · Kahn1 vs Jev on JevBench

JevBench is an independent public benchmark for Jev-class decision models, with long rubrics and hard judgment calls. Jev's per-item outcomes on the 231 public items come from JevBench's own results (v1.3.0). JevK5 v0.2, another open model (Qwen3.5-4B with a LoRA distilled from Qwen3.6-27B), is its own public run through JevBench's runner. Kahn1 answered the same items, k = 3, calibrated. The held-out half of JevBench is not public, so Kahn1 was not run on it.

TierPublic itemsKahn1 4BKahn1 3BJevK5 v0.2Jev 1.13.0Jev, public + held-out
Easy48100.0%100.0%100.0%100.0%100.0%
Standard7297.2%84.7%95.8%98.6%99.0%
Hard11175.7%42.3%73.9%73.0%74.1%
All23187.4%67.5%86.1%86.6%—

Kahn1 4B against JevK5's own published run, item by item: 13 items only Kahn1 4B gets right, 10 only JevK5; exact McNemar test, p = 0.68, not a significant difference. Against Jev's published outcomes: 13 items only Kahn1 4B gets right, 11 only Jev; p = 0.84, not a significant difference either. JevBench's own run of JevK5 v0.2 scores 85.3% on the same 231 items (v1.4 aggregates). The Kahn1, JevK5 and Jev figures come from three different runners on the same items: ours, JevK5's authors' and JevBench's. On the hard tier the 4B scores 75.7%, against 42.3% for the 3B and 73.0% for Jev; the gaps with Jev and JevK5 are not significant.

04 · Hard decisions on long documents

317 decision questions written by Claude Opus on long, realistic documents, in English and French, each checked by two blind Opus solvers. Kahn1 never trained on them, but they served to choose the 4B checkpoint: read them as a development split, not a benchmark. No other system was run on them. k = 1.

QuestionsItemsQwen3.5-4B, baseKahn1 4B
English23348.9%67.0%
French8448.8%69.0%
All31748.9%67.5%

The fine-tuning adds 19 points over the base model it starts from. Questions that need a date, a duration or an amount computed stay the weakest: one forward pass does not do arithmetic.

05 · Kahn1 in detail

Kahn1 4B on held-out data never seen in training, with k = 3 option permutations and temperature calibration, on a single RTX 5070 Ti (16 GB) with vLLM. The Kahn1 3B figures stay in the held-out table above.

Accuracy by dataset, Kahn1 4B

banking77 CHOICE · 8 OF 77 INTENTS · 3,076 ITEMS91.8%
MASSIVE CHOICE · 8 OF 60 INTENTS · 2,974 ITEMS92.9%
RTE NOUL · 277 ITEMS85.9%
SciTail NOUL · 2,126 ITEMS74.1%
SST-5 SCORE · 5 LEVELS · 2,210 ITEMS55.0%
App reviews SCORE · 5 LEVELS · 4,000 ITEMS50.2%

chance level (1/n options; each Choice question lists 8 of the intents). Score is exact match: on SST-5, human annotators agree only about 55 to 60% of the time.

By primitive, Kahn1 4B

PrimitiveItemsAccuracyECEBrier
Choice6,05092.35%0.0150.113
Noul2,40375.49%0.1270.346
Score6,21051.88%0.1120.624
All14,66372.45%0.0720.367

ECE: expected calibration error, 15 bins (0 is perfect). In an ablation on the 3B, an isotonic stage on top of temperature scaling raised overall ECE to 0.093, so Kahn1 uses temperature scaling only.

Ordinal scores, Kahn1 4B

DatasetExact±1 levelMAE E[S]Spearman ρ
SST-554.98%94.48%0.5720.833
App reviews50.18%84.42%0.6730.783

When the 4B misses the exact level it is usually one step away: 94.5% (SST-5) and 84.4% (app reviews) of its answers are within one level.

Snake: a decision every tick

20 games on a 10×10 board, the same rules and state text as the browser demo. Mean fruits eaten per game, bars scaled to 25.

Kahn1 3B, board + one-step readout MAX 40 · 22.2 MS PER DECISION, 20 GAMES BATCHED21.05
Greedy heuristic MAX 3421.00
Random move, never straight back MAX 10.20
Kahn1 3B, bare board 5 STEPS ON AVERAGE, THEN A WALL0.00

From the bare ASCII board it cannot do the geometry. Once the state says what lies one step in each direction, it plays as well as the heuristic.

Prefix caching

Median latency for N questions asked about one state, Kahn1 4B (bf16, vLLM, one RTX 5070 Ti, k = 1, median of 10 runs). Qwen3.5 is a hybrid model (Gated DeltaNet and attention layers), so vLLM caches its prefix in blocks of 528 tokens. On a 1,064-token state, 93% of each prompt comes from the cache and 10 questions take 143 ms, about 3× one question (about 10 ms per extra question). On a 212-token state no block fills, nothing is reused and each question recomputes the state: latency grows almost linearly, 319 ms for 10.

Questions1246810From the cache
212-token state45.5 ms83.8 ms136.4 ms189.4 ms249.1 ms318.8 ms0%
1,064-token state49.3 ms92.8 ms108.7 ms104.1 ms121.7 ms142.7 ms93%

Full reports: Kahn1 4B · held-out 4B · JevBench 4B · prefix caching · Kahn1 vs JevK5 · Choice over the same options (3B) · held-out 3B · JevBench 3B · Snake benchmark

06 · Latency and price

The latencies are not the same measurement: Kahn1 runs on a local GPU, JEV is a request over the internet.

Kahn1 4BKahn1 3BJEV 1.13.0
Whereyour hardware: vLLM on one RTX 5070 Ti (16 GB); also runs on a CPU, at seconds per questionthe same, on the same GPUTypeSafe's hosted API
Median latency88.7 ms, p95 279.3 ms (k = 3)36.4 ms, p95 82.6 ms (k = 3)248 ms, p95 308 ms, round trip from our client, network included, 6 requests in flight
Throughput8.5 queries/s (k = 3)21.7 queries/s, one at a time23.3 items/s with 6 in flight
Priceyour GPU or CPU; weights Apache 2.0your GPU or CPU; weights under the Qwen Research License (a research licence, see its terms)$0.042 per million input tokens, output free (public tariff)

07 · Reproduce

$ python scripts/fit_calibration.py --model Okura66/Kahn1-Qwen3.5-4B --format qwen3 --out calibration_4b.json
$ python -m eval.eval_full --model Okura66/Kahn1-Qwen3.5-4B --prompt-format qwen3 --eval data/eval.jsonl \
    --calibration calibration_4b.json --n-permutations 3 --max-model-len 4096     # Kahn1 4B, held-out
$ python scripts/jev_holdout.py run                              # JEV on data/eval.jsonl, resumable
$ python scripts/jev_holdout.py run --noul-template supported   # Noul items, reworded
$ python scripts/choice_fairness.py jev8                        # JEV, Choice over the same 8 options
$ python scripts/jevbench_vs_jev.py                             # reports/JEVBENCH_VS_JEV.md
$ python scripts/kahn1_4b_report.py                             # reports/KAHN1_4B_REPORT.md

JEV needs JEV_API_KEY. Build the held-out split first with python training/build_dataset.py --eval-only --eval-out data/eval.jsonl. Reports: Kahn1 4B vs 3B, JEV and JevK5 · Kahn1 4B held-out · Kahn1 4B JevBench · Kahn1 vs JevK5 · Choice over the same options (3B) · JevBench, Kahn1 3B vs Jev · Kahn1 3B held-out. JEV_VS_KAHN1.md is the first run, which gave JEV every intent on Choice.