Kahn1 vs Laya
Kahn1 4B, a 4B language model, against Laya, a 421M encoder from Convai Innovations: two open decision models that take Jev-style questions, on the same items, on the same GPU.
Updated October 2, 2026
01 · The short answer
Kahn1 4B is far ahead: 72.4% against 59.1% on our held-out set, although Laya was trained on its six sources, and 87.4% against 57.6% on JevBench (hard tier 75.7% against 32.4%). The gap is widest on Score (51.9% against 29.5%); on Noul the two are close (75.5% against 74.8%). Laya is ten times smaller, runs on a CPU and answers in about 24 ms on our GPU.
02 · Side by side
| Kahn1 4B | Laya | |
|---|---|---|
| Developer | Kahn1, independent | Convai Innovations |
| Licence | Weights Apache 2.0, code MIT | Apache 2.0 |
| Built on | Qwen3.5-4B + LoRA, merged | ModernBERT-large encoder (English), mmBERT-base (multilingual) |
| Size | 4B; 8.4 GB in bf16 | 421M (843 MB) and 322M (644 MB) |
| Inputs | Text | Text; 512 tokens (English), 1,024 to 8,192 (multilingual) |
| How answers are scored | Option-token logits, 3 option orders averaged, temperature calibration | A policy head trained with proper scoring rules; shipped temperatures |
| Request format | Jev question fields under a "schema" key at /v1/evaluate/jev (adapter needed for a Jev client) | Jev /v1/systemone request shape through laya-serve |
| Runs on | One 16 GB GPU in bf16 (vLLM), or a CPU (transformers) | A CPU or any GPU |
| Latency | Median 88.7 ms per request, k = 3, RTX 5070 Ti (our measure) | Median 23.5 ms per question on our RTX 5070 Ti (its Router, one question per request) |
03 · Measured on the same items
Kahn1 held-out set (14,663 items, Choice over the same 8 options):
| Group | Items | Kahn1 4B | Laya |
|---|---|---|---|
| All | 14,663 | 72.4% | 59.1% |
| Choice | 6,050 | 92.3% | 83.2% |
| Score | 6,210 | 51.9% | 29.5% |
| Noul | 2,403 | 75.5% | 74.8% |
| BANKING77 | 3,076 | 91.8% | 86.1% |
| MASSIVE | 2,974 | 92.9% | 80.1% |
| SST-5 | 2,210 | 55.0% | 40.4% |
| App reviews | 4,000 | 50.2% | 23.5% |
| SciTail | 2,126 | 74.1% | 74.5% |
| RTE | 277 | 85.9% | 77.3% |
Public JevBench (231 items):
| Tier | Items | Kahn1 4B | Laya |
|---|---|---|---|
| All | 231 | 87.4% | 57.6% |
| Standard | 72 | 97.2% | 70.8% |
| Easy | 48 | 100.0% | 95.8% |
| Hard | 111 | 75.7% | 32.4% |
Paired item by item (exact McNemar test): on the held-out set, Laya gets 1,010 items right that Kahn1 misses and Kahn1 2,969 that Laya misses (p = 6 × 10−221); on JevBench, 8 against 77 (p = 3 × 10−15). ECE: Kahn1 4B 0.072, Laya 0.184. Full method and biases: open decision models, measured like for like.
04 · Which one to pick
Pick Laya if
- You need a CPU-only model, or the lowest latency.
- Your texts are short and your tasks close to its training sets (intents, sentiment, entailment).
- You want Jev's request shape through laya-serve.
Pick Kahn1 if
- You ask for ratings on a scale (Score): Laya mostly answers the lowest levels on app reviews.
- Your texts are longer than 512 tokens, or the decisions are hard (JevBench hard tier).
- You threshold on the confidence: Kahn1 is the better calibrated (ECE 0.072 against 0.184).
05 · What to keep in mind
- Laya's held-out figures are in-distribution: its model card lists the six held-out sources among its training data.
- Laya's English checkpoint truncates 57 of the 231 JevBench states (512 tokens); its multilingual checkpoint, reading them in full, scores lower.
- Laya answers in one pass with its shipped calibration; Kahn1 averages 3 option orders and calibrates.
- Latency: Laya's 23.5 ms is one question per request through its Router on our GPU; Kahn1's 88.7 ms is k = 3 through vLLM. Different settings.
06 · FAQ
Is Laya more accurate than Kahn1?
No. On the same 14,663 held-out items Laya scores 59.1% and Kahn1 4B 72.4%; on the 231 public JevBench items, 57.6% against 87.4%.
Is Laya faster?
Yes: a 421M encoder answers in about 24 ms per question on our GPU, and runs on a CPU. Kahn1 4B takes about 89 ms at k = 3 on the same GPU.
Was Laya tested on data it trained on?
Its model card lists BANKING77, MASSIVE, SST-5, SciTail, RTE and app_reviews among its training data, the six sources of Kahn1's held-out set, so its held-out figures are in-distribution. JevBench is not among them.