Benchmarks
Every number behind the home page: Kahn1 4B and Kahn1 3B against JEV on the same items, Kahn1's own held-out results, latency, and how to reproduce them.
01 · At a glance
Two benchmarks, both paired item by item. On the Kahn1 held-out set JEV was run by us through its API, with the same states, prompts and options as Kahn1. On JevBench, Jev's results are the ones JevBench itself publishes. The chart shows Kahn1 4B and Kahn1 3B; the tables below add the per-source figures, and JevK5 on JevBench.
Other open models. Run by us on the same items and the same GPU, Cloudflare's Clef-flash (9B, in int8) is the most accurate on the held-out set, 74.8% against 72.4% for Kahn1 4B and 73.2% for JEV, and level with Kahn1 4B on JevBench (83.5%, not a significant gap); Laya (421M) scores 59.1% and 57.6%. Every table, the test method and the known biases: open decision models, measured like for like.
Correction. An earlier version of this page asked Choice over 8 options for Kahn1 but over every intent for JEV, which made Kahn1 look ahead on Choice and overall. That comparison was unequal and has been corrected: over the same options, JEV is ahead on Choice and overall. Both conditions, side by side (3B).
TABLE VIEW
| Benchmark | Group | Items | Kahn1 4B | Kahn1 3B | JEV |
|---|---|---|---|---|---|
| KAHN1 HELD-OUT | Choice | 6,050 | 92.3% | 90.9% | 94.5% |
| KAHN1 HELD-OUT | Score | 6,210 | 51.9% | 51.4% | 52.0% |
| KAHN1 HELD-OUT | Noul | 2,403 | 75.5% | 67.0% | 74.4% |
| KAHN1 HELD-OUT | All | 14,663 | 72.4% | 70.3% | 73.2% |
| JEVBENCH PUBLIC | Easy | 48 | 100.0% | 100.0% | 100.0% |
| JEVBENCH PUBLIC | Standard | 72 | 97.2% | 84.7% | 98.6% |
| JEVBENCH PUBLIC | Hard | 111 | 75.7% | 42.3% | 73.0% |
| JEVBENCH PUBLIC | All | 231 | 87.4% | 67.5% | 86.6% |
02 · Kahn1 vs JEV on the Kahn1 held-out set
14,663 items from six public datasets, none seen in training, in the task families Kahn1 was trained on: intents (banking77, MASSIVE), 5-level scales (SST-5, app reviews) and entailment (RTE, SciTail). One question per request, and the same options for every system: each Choice question lists the right intent and 7 distractors seeded from the state, 8 options in all. Kahn1 4B and Kahn1 3B: k = 3, temperature calibration. JEV: jev-1.13.0, answered all 14,663 without error. ECE: 15 bins on pmax, 0 is perfect.
| Dataset | Items | Kahn1 4B | Kahn1 3B | JEV | 4B ECE | 3B ECE | JEV ECE |
|---|---|---|---|---|---|---|---|
| banking77 | 3,076 | 91.8% | 91.1% | 94.8% | 0.018 | 0.026 | 0.024 |
| MASSIVE | 2,974 | 92.9% | 90.8% | 94.1% | 0.013 | 0.011 | 0.017 |
| RTE | 277 | 85.9% | 82.7% | 89.9% | 0.048 | 0.105 | 0.034 |
| SciTail | 2,126 | 74.1% | 65.0% | 72.4% | 0.141 | 0.248 | 0.077 |
| SST-5 | 2,210 | 55.0% | 52.4% | 57.7% | 0.052 | 0.098 | 0.178 |
| App reviews | 4,000 | 50.2% | 50.9% | 48.9% | 0.146 | 0.096 | 0.252 |
| Choice | 6,050 | 92.3% | 90.9% | 94.5% | 0.015 | 0.017 | 0.019 |
| Score | 6,210 | 51.9% | 51.4% | 52.0% | 0.112 | 0.097 | 0.226 |
| Noul | 2,403 | 75.5% | 67.0% | 74.4% | 0.127 | 0.232 | 0.071 |
| All | 14,663 | 72.4% | 70.3% | 73.2% | 0.072 | 0.082 | 0.113 |
JEV is ahead on Choice and overall and level on Score (52.0% against 51.9%); Kahn1 4B is ahead on Noul (75.5% against 74.4%, a gap not tested for significance). The 4B is ahead of the 3B on every primitive and the better calibrated overall.
Choice over every intent
The table above asks Choice over 8 options. Asked over every intent instead (77 for banking77, 60 for MASSIVE), on the same 1,184 items, every system drops, Kahn1 more than JEV. Kahn1 goes through its two-stage router there: one Noul per option keeps the 10 best, then a Choice picks among them (k = 1).
| Options per question | Items | Kahn1 4B | Kahn1 3B | JEV |
|---|---|---|---|---|
| 8 (the right one + 7 seeded distractors) | 6,050 | 92.3% | 90.9% | 94.5% |
| every intent (77 / 60) | 1,184 | 71.6% | 68.4% | 79.1% |
How the Noul question is worded
JEV's native Noul form is the bare statement as instructions, which is also what JevBench sends.
JEV then judges whether the statement is true in general, not whether the text supports it: on SciTail, whose
negatives are statements the text does not support, it answers yes to 93% of items. Kahn1's prompt asks whether
the statement is "true for this state", so the Noul items were re-run with
The text supports this statement: {statement}. The tables above use that wording, the fairer one for JEV.
| Dataset | Items | Kahn1 4B | Kahn1 3B | JEV, bare statement | JEV, “supported” wording |
|---|---|---|---|---|---|
| RTE | 277 | 85.9% | 82.7% | 82.3% | 89.9% |
| SciTail | 2,126 | 74.1% | 65.0% | 44.4% | 72.4% |
| Noul | 2,403 | 75.5% | 67.0% | 48.7% | 74.4% |
03 · Kahn1 vs Jev on JevBench
JevBench is an independent public benchmark for Jev-class decision models, with long rubrics and hard judgment calls. Jev's per-item outcomes on the 231 public items come from JevBench's own results (v1.3.0). JevK5 v0.2, another open model (Qwen3.5-4B with a LoRA distilled from Qwen3.6-27B), is its own public run through JevBench's runner. Kahn1 answered the same items, k = 3, calibrated. The held-out half of JevBench is not public, so Kahn1 was not run on it.
| Tier | Public items | Kahn1 4B | Kahn1 3B | JevK5 v0.2 | Jev 1.13.0 | Jev, public + held-out |
|---|---|---|---|---|---|---|
| Easy | 48 | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% |
| Standard | 72 | 97.2% | 84.7% | 95.8% | 98.6% | 99.0% |
| Hard | 111 | 75.7% | 42.3% | 73.9% | 73.0% | 74.1% |
| All | 231 | 87.4% | 67.5% | 86.1% | 86.6% | — |
Kahn1 4B against JevK5's own published run, item by item: 13 items only Kahn1 4B gets right, 10 only JevK5; exact McNemar test, p = 0.68, not a significant difference. Against Jev's published outcomes: 13 items only Kahn1 4B gets right, 11 only Jev; p = 0.84, not a significant difference either. JevBench's own run of JevK5 v0.2 scores 85.3% on the same 231 items (v1.4 aggregates). The Kahn1, JevK5 and Jev figures come from three different runners on the same items: ours, JevK5's authors' and JevBench's. On the hard tier the 4B scores 75.7%, against 42.3% for the 3B and 73.0% for Jev; the gaps with Jev and JevK5 are not significant.
04 · Hard decisions on long documents
317 decision questions written by Claude Opus on long, realistic documents, in English and French, each checked by two blind Opus solvers. Kahn1 never trained on them, but they served to choose the 4B checkpoint: read them as a development split, not a benchmark. No other system was run on them. k = 1.
| Questions | Items | Qwen3.5-4B, base | Kahn1 4B |
|---|---|---|---|
| English | 233 | 48.9% | 67.0% |
| French | 84 | 48.8% | 69.0% |
| All | 317 | 48.9% | 67.5% |
The fine-tuning adds 19 points over the base model it starts from. Questions that need a date, a duration or an amount computed stay the weakest: one forward pass does not do arithmetic.
05 · Kahn1 in detail
Kahn1 4B on held-out data never seen in training, with k = 3 option permutations and temperature calibration, on a single RTX 5070 Ti (16 GB) with vLLM. The Kahn1 3B figures stay in the held-out table above.
Accuracy by dataset, Kahn1 4B
chance level (1/n options; each Choice question lists 8 of the intents). Score is exact match: on SST-5, human annotators agree only about 55 to 60% of the time.
By primitive, Kahn1 4B
| Primitive | Items | Accuracy | ECE | Brier |
|---|---|---|---|---|
| Choice | 6,050 | 92.35% | 0.015 | 0.113 |
| Noul | 2,403 | 75.49% | 0.127 | 0.346 |
| Score | 6,210 | 51.88% | 0.112 | 0.624 |
| All | 14,663 | 72.45% | 0.072 | 0.367 |
ECE: expected calibration error, 15 bins (0 is perfect). In an ablation on the 3B, an isotonic stage on top of temperature scaling raised overall ECE to 0.093, so Kahn1 uses temperature scaling only.
Ordinal scores, Kahn1 4B
| Dataset | Exact | ±1 level | MAE E[S] | Spearman ρ |
|---|---|---|---|---|
| SST-5 | 54.98% | 94.48% | 0.572 | 0.833 |
| App reviews | 50.18% | 84.42% | 0.673 | 0.783 |
When the 4B misses the exact level it is usually one step away: 94.5% (SST-5) and 84.4% (app reviews) of its answers are within one level.
Snake: a decision every tick
20 games on a 10×10 board, the same rules and state text as the browser demo. Mean fruits eaten per game, bars scaled to 25.
From the bare ASCII board it cannot do the geometry. Once the state says what lies one step in each direction, it plays as well as the heuristic.
Prefix caching
Median latency for N questions asked about one state, Kahn1 4B (bf16, vLLM, one RTX 5070 Ti, k = 1, median of 10 runs). Qwen3.5 is a hybrid model (Gated DeltaNet and attention layers), so vLLM caches its prefix in blocks of 528 tokens. On a 1,064-token state, 93% of each prompt comes from the cache and 10 questions take 143 ms, about 3× one question (about 10 ms per extra question). On a 212-token state no block fills, nothing is reused and each question recomputes the state: latency grows almost linearly, 319 ms for 10.
| Questions | 1 | 2 | 4 | 6 | 8 | 10 | From the cache |
|---|---|---|---|---|---|---|---|
| 212-token state | 45.5 ms | 83.8 ms | 136.4 ms | 189.4 ms | 249.1 ms | 318.8 ms | 0% |
| 1,064-token state | 49.3 ms | 92.8 ms | 108.7 ms | 104.1 ms | 121.7 ms | 142.7 ms | 93% |
Full reports: Kahn1 4B · held-out 4B · JevBench 4B · prefix caching · Kahn1 vs JevK5 · Choice over the same options (3B) · held-out 3B · JevBench 3B · Snake benchmark
06 · Latency and price
The latencies are not the same measurement: Kahn1 runs on a local GPU, JEV is a request over the internet.
| Kahn1 4B | Kahn1 3B | JEV 1.13.0 | |
|---|---|---|---|
| Where | your hardware: vLLM on one RTX 5070 Ti (16 GB); also runs on a CPU, at seconds per question | the same, on the same GPU | TypeSafe's hosted API |
| Median latency | 88.7 ms, p95 279.3 ms (k = 3) | 36.4 ms, p95 82.6 ms (k = 3) | 248 ms, p95 308 ms, round trip from our client, network included, 6 requests in flight |
| Throughput | 8.5 queries/s (k = 3) | 21.7 queries/s, one at a time | 23.3 items/s with 6 in flight |
| Price | your GPU or CPU; weights Apache 2.0 | your GPU or CPU; weights under the Qwen Research License (a research licence, see its terms) | $0.042 per million input tokens, output free (public tariff) |
07 · Reproduce
$ python scripts/fit_calibration.py --model Okura66/Kahn1-Qwen3.5-4B --format qwen3 --out calibration_4b.json $ python -m eval.eval_full --model Okura66/Kahn1-Qwen3.5-4B --prompt-format qwen3 --eval data/eval.jsonl \ --calibration calibration_4b.json --n-permutations 3 --max-model-len 4096 # Kahn1 4B, held-out $ python scripts/jev_holdout.py run # JEV on data/eval.jsonl, resumable $ python scripts/jev_holdout.py run --noul-template supported # Noul items, reworded $ python scripts/choice_fairness.py jev8 # JEV, Choice over the same 8 options $ python scripts/jevbench_vs_jev.py # reports/JEVBENCH_VS_JEV.md $ python scripts/kahn1_4b_report.py # reports/KAHN1_4B_REPORT.md
JEV needs JEV_API_KEY. Build the held-out split first with
python training/build_dataset.py --eval-only --eval-out data/eval.jsonl. Reports:
Kahn1 4B vs 3B, JEV and JevK5 ·
Kahn1 4B held-out ·
Kahn1 4B JevBench ·
Kahn1 vs JevK5 ·
Choice over the same options (3B) ·
JevBench, Kahn1 3B vs Jev ·
Kahn1 3B held-out.
JEV_VS_KAHN1.md is the first run, which gave JEV every intent on
Choice.