KAHN1
EN FR

Kahn1 vs Clef-flash

Two open decision models that answer typed questions about a text with a probability per option: Kahn1 4B and Cloudflare's Clef-flash, on the same items, run on the same GPU.

Updated October 2, 2026

01 · The short answer

Clef-flash is more accurate on our held-out set (74.8% against 72.4%, a significant gap), above all on Choice (98.6% against 92.3%). Kahn1 4B is ahead on Noul (75.5% against 69.7%), better calibrated (ECE 0.072 against 0.101) and level on JevBench (87.4% against 83.5%, not significant; hard tier 75.7% against 67.6%). Kahn1 4B is half the size and runs in bf16 on a 16 GB card; Clef-flash also reads images and video, and speaks Jev's request format as it is.

02 · Side by side

Kahn1 4BClef-flash
DeveloperKahn1, independentCloudflare
LicenceWeights Apache 2.0, code MITApache 2.0
Built onQwen3.5-4B + LoRA, mergedQwen3.5-9B + a joint schema head
Size4B; 8.4 GB in bf169B; 18.8 GB in bf16
InputsTextText, JSON, images, video; up to 16,384 tokens by default
How answers are scoredOption-token logits, 3 option orders averaged, temperature calibrationA head that scores every option of every question jointly; calibration according to the vendor
Request formatJev question fields under a "schema" key at /v1/evaluate/jev (adapter needed for a Jev client)Jev /v1/systemone request and response body (release code), hosted on Workers AI
Runs onOne 16 GB GPU in bf16 (vLLM), or a CPU (transformers)A GPU with more than 16 GB in bf16 (vendor tested an H200); we ran it in int8 on 16 GB
LatencyMedian 88.7 ms per request, k = 3, RTX 5070 Ti (our measure)Median 38.8 ms per request according to the vendor; not measured comparably here

03 · Measured on the same items

Kahn1 held-out set (14,663 items, Choice over the same 8 options):

GroupItemsKahn1 4BClef-flash
All14,66372.4%74.8%
Choice6,05092.3%98.6%
Score6,21051.9%53.7%
Noul2,40375.5%69.7%
BANKING773,07691.8%99.3%
MASSIVE2,97492.9%97.8%
SST-52,21055.0%58.2%
App reviews4,00050.2%51.2%
SciTail2,12674.1%67.6%
RTE27785.9%85.9%

Public JevBench (231 items):

TierItemsKahn1 4BClef-flash
All23187.4%83.5%
Standard7297.2%97.2%
Easy48100.0%100.0%
Hard11175.7%67.6%

Paired item by item (exact McNemar test): on the held-out set, Clef-flash gets 1,187 items right that Kahn1 misses and Kahn1 836 that Clef-flash misses (p = 6 × 10−15); on JevBench, 12 against 21 (p = 0.16). ECE: Kahn1 4B 0.072, Clef-flash 0.101. Full method and biases: open decision models, measured like for like.

04 · Which one to pick

Pick Clef-flash if

  • You want the best measured accuracy on intent-style Choice questions.
  • Your states include images, video or long JSON.
  • You want Jev's request and response format with no adapter, or Cloudflare's hosting (Workers AI).
  • You have a GPU with more than 16 GB, or accept int8 on 16 GB.

Pick Kahn1 if

  • You threshold on the confidence: Kahn1 is the better calibrated of the two.
  • Your questions are mostly entailment and policy checks (Noul).
  • You have one 16 GB GPU, or only a CPU.
  • You want the smaller model at the same JevBench level.

05 · What to keep in mind

  • Clef-flash's training data is not published. Our train-split probe finds no memorisation, which does not rule out exposure, and Clef-flash removes 92% of its base model's errors on BANKING77 where Kahn1 removes 27% of its own: its Choice lead may not be zero-shot (see the method page).
  • Clef-flash ran in int8 (bitsandbytes), not in bf16 as released.
  • Kahn1 averages 3 option orders and calibrates; Clef-flash answers in one pass, its probabilities as shipped.
  • The held-out set is Kahn1's; JevBench is independent but only its public half was run.

06 · FAQ

Is Clef-flash more accurate than Kahn1?

On Kahn1's held-out set, yes: 74.8% against 72.4% on the same 14,663 items, a significant gap, driven by Choice. On the 231 public JevBench items, the two are level (87.4% for Kahn1 4B, 83.5% for Clef-flash, exact McNemar p = 0.16).

Can Clef-flash run on a 16 GB GPU?

Not in bf16: its weights take 18.8 GB. We ran it in int8 with bitsandbytes on a 16 GB RTX 5070 Ti; that changes the model slightly.

Which is better calibrated?

Kahn1 4B: ECE 0.072 against 0.101 for Clef-flash on the held-out set, each as shipped.

Did Clef-flash see the test data?

We cannot tell. Its training data is not published. A probe on the train splits of the test sources finds no memorisation, but a model trained without overfitting would not show any either. Against its own base model (Qwen3.5-9B), Clef-flash removes 92% of the errors on BANKING77; Kahn1 4B, which never saw it, removes 27% of its base model's.