Models
Kahn1 comes in two sizes, Kahn1 4B and Kahn1 3B: the same model, trained for the same three primitives and served by the same engine. Each is on Hugging Face as a merged model and as a LoRA adapter.
01 · The two sizes
| Kahn1 4B | Kahn1 3B | |
|---|---|---|
| Base model | Qwen/Qwen3.5-4B | Qwen/Qwen2.5-3B-Instruct |
| Merged model | Okura66/Kahn1-Qwen3.5-4B · 8.4 GB | Okura66/Kahn1-Qwen2.5-3B · 6.17 GB |
| LoRA adapter | Okura66/Kahn1-Qwen3.5-4B-LoRA · 57 MB | Okura66/Kahn1-Qwen2.5-3B-LoRA · 239 MB |
| Prompt format | native Qwen3 chat template, picked automatically | tag layout, picked automatically |
| Licence | Apache 2.0 | Qwen Research License, a research licence: see its terms |
| Held-out, 14,663 items | 72.4% | 70.3% |
| JevBench, 231 public items | 87.4% (hard 75.7%) | 67.5% (hard 42.3%) |
| Median latency, k = 3 | 88.7 ms | 36.4 ms |
| Calibration | calibration.json in the Hugging Face repository | calibration.json in the Hugging Face repository |
Held-out and JevBench: k = 3, temperature calibration. Latency: vLLM on one RTX 5070 Ti. Details on the benchmarks page. The code is MIT.
GGUF quantizations: mradermacher/Kahn1-Qwen2.5-3B-GGUF,
a smaller build of the 3B by mradermacher, older than the published 3B weights. They are what the browser demos
(playground, Snake) load.
02 · Which one
Both sizes run on a GPU (vLLM) and on a CPU (transformers), with the same API and the same answer format. The facts: on the held-out set the 4B scores 72.4% and the 3B 70.3%; on JevBench the 4B scores 87.4% and the 3B 67.5%, and 75.7% against 42.3% on the hard tier. The 3B is smaller (6.17 GB against 8.4 GB) and faster (median 36.4 ms against 88.7 ms). The 4B's weights are Apache 2.0, the 3B's are under the Qwen Research License.
03 · Load one
With sysone, set the model id; the engine picks its prompt format by itself.
$ SYSONE_MODEL=Okura66/Kahn1-Qwen3.5-4B uv run sysone serve --port 8000 # GPU, vLLM $ SYSONE_BACKEND=cpu SYSONE_MODEL=Okura66/Kahn1-Qwen3.5-4B uv run sysone serve --port 8000 # CPU # Kahn1 3B: SYSONE_MODEL=Okura66/Kahn1-Qwen2.5-3B, on either backend
A LoRA adapter goes on top of its base model with peft:
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B", dtype=torch.bfloat16)
tokenizer = AutoTokenizer.from_pretrained("Okura66/Kahn1-Qwen3.5-4B-LoRA")
model = PeftModel.from_pretrained(base, "Okura66/Kahn1-Qwen3.5-4B-LoRA")
# Kahn1 3B: base "Qwen/Qwen2.5-3B-Instruct", adapter "Okura66/Kahn1-Qwen2.5-3B-LoRA"
To serve an adapter with vLLM, merge it first:
$ python scripts/merge_qwen_lora.py --base Qwen/Qwen3.5-4B --adapter Okura66/Kahn1-Qwen3.5-4B-LoRA --output checkpoints/kahn1_4b
The guide covers the API, the answer format and calibration.
Hugging Face cards: Kahn1 4B · Kahn1 4B LoRA · Kahn1 3B · Kahn1 3B LoRA · all benchmark results