Kahn1 vs Clef-flash
Two open decision models that answer typed questions about a text with a probability per option: Kahn1 4B and Cloudflare's Clef-flash, on the same items, run on the same GPU.
Updated October 2, 2026
01 · The short answer
Clef-flash is more accurate on our held-out set (74.8% against 72.4%, a significant gap), above all on Choice (98.6% against 92.3%). Kahn1 4B is ahead on Noul (75.5% against 69.7%), better calibrated (ECE 0.072 against 0.101) and level on JevBench (87.4% against 83.5%, not significant; hard tier 75.7% against 67.6%). Kahn1 4B is half the size and runs in bf16 on a 16 GB card; Clef-flash also reads images and video, and speaks Jev's request format as it is.
02 · Side by side
| Kahn1 4B | Clef-flash | |
|---|---|---|
| Developer | Kahn1, independent | Cloudflare |
| Licence | Weights Apache 2.0, code MIT | Apache 2.0 |
| Built on | Qwen3.5-4B + LoRA, merged | Qwen3.5-9B + a joint schema head |
| Size | 4B; 8.4 GB in bf16 | 9B; 18.8 GB in bf16 |
| Inputs | Text | Text, JSON, images, video; up to 16,384 tokens by default |
| How answers are scored | Option-token logits, 3 option orders averaged, temperature calibration | A head that scores every option of every question jointly; calibration according to the vendor |
| Request format | Jev question fields under a "schema" key at /v1/evaluate/jev (adapter needed for a Jev client) | Jev /v1/systemone request and response body (release code), hosted on Workers AI |
| Runs on | One 16 GB GPU in bf16 (vLLM), or a CPU (transformers) | A GPU with more than 16 GB in bf16 (vendor tested an H200); we ran it in int8 on 16 GB |
| Latency | Median 88.7 ms per request, k = 3, RTX 5070 Ti (our measure) | Median 38.8 ms per request according to the vendor; not measured comparably here |
03 · Measured on the same items
Kahn1 held-out set (14,663 items, Choice over the same 8 options):
| Group | Items | Kahn1 4B | Clef-flash |
|---|---|---|---|
| All | 14,663 | 72.4% | 74.8% |
| Choice | 6,050 | 92.3% | 98.6% |
| Score | 6,210 | 51.9% | 53.7% |
| Noul | 2,403 | 75.5% | 69.7% |
| BANKING77 | 3,076 | 91.8% | 99.3% |
| MASSIVE | 2,974 | 92.9% | 97.8% |
| SST-5 | 2,210 | 55.0% | 58.2% |
| App reviews | 4,000 | 50.2% | 51.2% |
| SciTail | 2,126 | 74.1% | 67.6% |
| RTE | 277 | 85.9% | 85.9% |
Public JevBench (231 items):
| Tier | Items | Kahn1 4B | Clef-flash |
|---|---|---|---|
| All | 231 | 87.4% | 83.5% |
| Standard | 72 | 97.2% | 97.2% |
| Easy | 48 | 100.0% | 100.0% |
| Hard | 111 | 75.7% | 67.6% |
Paired item by item (exact McNemar test): on the held-out set, Clef-flash gets 1,187 items right that Kahn1 misses and Kahn1 836 that Clef-flash misses (p = 6 × 10−15); on JevBench, 12 against 21 (p = 0.16). ECE: Kahn1 4B 0.072, Clef-flash 0.101. Full method and biases: open decision models, measured like for like.
04 · Which one to pick
Pick Clef-flash if
- You want the best measured accuracy on intent-style Choice questions.
- Your states include images, video or long JSON.
- You want Jev's request and response format with no adapter, or Cloudflare's hosting (Workers AI).
- You have a GPU with more than 16 GB, or accept int8 on 16 GB.
Pick Kahn1 if
- You threshold on the confidence: Kahn1 is the better calibrated of the two.
- Your questions are mostly entailment and policy checks (Noul).
- You have one 16 GB GPU, or only a CPU.
- You want the smaller model at the same JevBench level.
05 · What to keep in mind
- Clef-flash's training data is not published. Our train-split probe finds no memorisation, which does not rule out exposure, and Clef-flash removes 92% of its base model's errors on BANKING77 where Kahn1 removes 27% of its own: its Choice lead may not be zero-shot (see the method page).
- Clef-flash ran in int8 (bitsandbytes), not in bf16 as released.
- Kahn1 averages 3 option orders and calibrates; Clef-flash answers in one pass, its probabilities as shipped.
- The held-out set is Kahn1's; JevBench is independent but only its public half was run.
06 · FAQ
Is Clef-flash more accurate than Kahn1?
On Kahn1's held-out set, yes: 74.8% against 72.4% on the same 14,663 items, a significant gap, driven by Choice. On the 231 public JevBench items, the two are level (87.4% for Kahn1 4B, 83.5% for Clef-flash, exact McNemar p = 0.16).
Can Clef-flash run on a 16 GB GPU?
Not in bf16: its weights take 18.8 GB. We ran it in int8 with bitsandbytes on a 16 GB RTX 5070 Ti; that changes the model slightly.
Which is better calibrated?
Kahn1 4B: ECE 0.072 against 0.101 for Clef-flash on the held-out set, each as shipped.
Did Clef-flash see the test data?
We cannot tell. Its training data is not published. A probe on the train splits of the test sources finds no memorisation, but a model trained without overfitting would not show any either. Against its own base model (Qwen3.5-9B), Clef-flash removes 92% of the errors on BANKING77; Kahn1 4B, which never saw it, removes 27% of its base model's.