KAHN1
EN FR

Caveats

Kahn1 comes in two sizes, 4 and 3 billion parameters. Read this before putting either in front of real decisions.

It is not a frontier model.

On subjective or ambiguous judgments it is wrong more often than a large model. On the hard tier of the external JevBench, Kahn1 4B is level with Jev (75.7% against 73.0%, not a significant gap) and the 3B gets only 42.3%. JEV is ahead on the held-out set overall and on Choice. Computing a date, a duration or an amount in a single forward pass is a weak spot.

Accuracy depends on the primitive.

Choice is strong (92% for the 4B, 91% for the 3B on held-out intent sets, 8 options per question). Score is the weakest primitive: 52% (4B) and 51% (3B) exact match on 5-level scales, though both are within one level about 95% of the time on SST-5. Noul: 86% on RTE and 74% on SciTail (out-of-domain science entailment) for the 4B, 83% and 65% for the 3B.

Calibration holds on its own distribution.

Temperatures were fitted on a calibration split drawn from the training distribution and kept out of training, not on your data, and Noul's calibration error is still 0.13 for the 4B, 0.23 for the 3B. Recalibrate on your domain before trusting a threshold: sysone calibrate, described in the guide.

Confidence is not the probability of being right.

Confidence is (p_max − 1/n) / (1 − 1/n): it measures the distance from a uniform guess. Set thresholds on held-out data from your domain and check the risk-coverage curve.

Deterministic computation is not deterministic truth.

Same weights, same prompt and same kernels give the same logits. That does not make the answer correct: adversarial or out-of-distribution text can still fool it, and it offers no formal guarantee.

Regulation still applies.

Reading logits instead of parsing JSON removes format errors, not the obligations that come with using a neural model for decisions (EU AI Act, sector rules). Keep a person or a larger model on the low-confidence path, and log what was decided.

The speed comes from engineering.

A single output position per question explains the latency, not a new architecture. It still grows with each question, almost linearly on short states, which vLLM cannot reuse for Kahn1 4B; long states are cached. On a CPU expect seconds per question; in a browser without WebGPU, minutes.

The browser demos run an older checkpoint.

The playground and Snake load a smaller 4-bit GGUF build of the 3B (thanks to mradermacher), older than the published 3B weights; the page flips its Noul answers, which that build learned inverted. The benchmarks are the unquantized Kahn1 4B and Kahn1 3B on vLLM.