KAHN1
EN FR

Models

Kahn1 comes in two sizes, Kahn1 4B and Kahn1 3B: the same model, trained for the same three primitives and served by the same engine. Each is on Hugging Face as a merged model and as a LoRA adapter.

01 · The two sizes

Kahn1 4BKahn1 3B
Base modelQwen/Qwen3.5-4BQwen/Qwen2.5-3B-Instruct
Merged modelOkura66/Kahn1-Qwen3.5-4B · 8.4 GBOkura66/Kahn1-Qwen2.5-3B · 6.17 GB
LoRA adapterOkura66/Kahn1-Qwen3.5-4B-LoRA · 57 MBOkura66/Kahn1-Qwen2.5-3B-LoRA · 239 MB
Prompt formatnative Qwen3 chat template, picked automaticallytag layout, picked automatically
LicenceApache 2.0Qwen Research License, a research licence: see its terms
Held-out, 14,663 items72.4%70.3%
JevBench, 231 public items87.4% (hard 75.7%)67.5% (hard 42.3%)
Median latency, k = 388.7 ms36.4 ms
Calibrationcalibration.json in the Hugging Face repositorycalibration.json in the Hugging Face repository

Held-out and JevBench: k = 3, temperature calibration. Latency: vLLM on one RTX 5070 Ti. Details on the benchmarks page. The code is MIT.

GGUF quantizations: mradermacher/Kahn1-Qwen2.5-3B-GGUF, a smaller build of the 3B by mradermacher, older than the published 3B weights. They are what the browser demos (playground, Snake) load.

02 · Which one

Both sizes run on a GPU (vLLM) and on a CPU (transformers), with the same API and the same answer format. The facts: on the held-out set the 4B scores 72.4% and the 3B 70.3%; on JevBench the 4B scores 87.4% and the 3B 67.5%, and 75.7% against 42.3% on the hard tier. The 3B is smaller (6.17 GB against 8.4 GB) and faster (median 36.4 ms against 88.7 ms). The 4B's weights are Apache 2.0, the 3B's are under the Qwen Research License.

03 · Load one

With sysone, set the model id; the engine picks its prompt format by itself.

$ SYSONE_MODEL=Okura66/Kahn1-Qwen3.5-4B uv run sysone serve --port 8000                      # GPU, vLLM
$ SYSONE_BACKEND=cpu SYSONE_MODEL=Okura66/Kahn1-Qwen3.5-4B uv run sysone serve --port 8000   # CPU
# Kahn1 3B: SYSONE_MODEL=Okura66/Kahn1-Qwen2.5-3B, on either backend

A LoRA adapter goes on top of its base model with peft:

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B", dtype=torch.bfloat16)
tokenizer = AutoTokenizer.from_pretrained("Okura66/Kahn1-Qwen3.5-4B-LoRA")
model = PeftModel.from_pretrained(base, "Okura66/Kahn1-Qwen3.5-4B-LoRA")
# Kahn1 3B: base "Qwen/Qwen2.5-3B-Instruct", adapter "Okura66/Kahn1-Qwen2.5-3B-LoRA"

To serve an adapter with vLLM, merge it first:

$ python scripts/merge_qwen_lora.py --base Qwen/Qwen3.5-4B --adapter Okura66/Kahn1-Qwen3.5-4B-LoRA --output checkpoints/kahn1_4b

The guide covers the API, the answer format and calibration.

Hugging Face cards: Kahn1 4B · Kahn1 4B LoRA · Kahn1 3B · Kahn1 3B LoRA · all benchmark results