Get started
From a clone to a calibrated answer: install Kahn1, serve it, send typed questions, read what comes back and tune it to your data. Everything here runs on your own machine.
01 · Let an AI agent set it up
Paste this prompt into a coding agent that can run commands on your machine (Claude Code, Codex, Cursor, Gemini CLI…): it checks your hardware, picks the GPU or CPU backend, installs, starts the server and sends a test request. It asks before sudo or system packages. A chat assistant without a terminal (ChatGPT or Claude in the browser) can walk you through the same steps, but cannot run them.
Set up Kahn1 on this machine and prove it works with one real request.
Kahn1 (https://github.com/Okura66/kahn1, MIT code) is an open-source decision engine: a fine-tuned open model that answers typed questions (choice, score, noul) about a text by reading the probabilities of option tokens. It comes in two sizes, and both run on a GPU (vLLM) or a CPU (transformers): Kahn1 4B (Okura66/Kahn1-Qwen3.5-4B, Apache 2.0 weights, 8.4 GB) and Kahn1 3B (Okura66/Kahn1-Qwen2.5-3B, weights under the Qwen Research License, a research licence: see its terms, 6.17 GB). The 3B is smaller and faster (median 36 ms against 89 ms on one GPU); the 4B is much stronger on hard judgment calls. Use the 4B unless I ask for the 3B. It is served by a FastAPI app, `sysone`. Guide: https://kahn1.com/get-started/
Rules: work step by step and show each command before you run it. Ask me before using sudo, installing system packages, or changing anything outside the project folder. If a step fails, show the exact error, explain the likely cause and propose a fix before going on. Never invent output: if the model could not run, say so.
1. Inspect the machine and tell me what you find: OS (Linux, macOS, Windows, WSL2), Python version (3.11+ needed), free RAM and disk space, and whether an NVIDIA GPU is usable (`nvidia-smi`). Then pick the backend and say why:
- GPU (vLLM): only on Linux or WSL2 with an NVIDIA GPU and CUDA. Kahn1 4B needs 12 GB of VRAM or more (it ran on 16 GB), Kahn1 3B 8 GB or more.
- CPU (transformers) otherwise: a few seconds per question. In float32, about 13 GB of free RAM for the 3B, about 17 GB for the 4B.
The download is 8.4 GB for the 4B, 6.17 GB for the 3B. If the chosen size does not fit this machine, tell me before going on.
2. Install uv if it is missing (https://docs.astral.sh/uv/). If this folder is not already a clone of the repository, clone it:
git clone https://github.com/Okura66/kahn1 && cd kahn1
uv venv --python 3.11
GPU: uv pip install -e ".[gpu]"
CPU: uv pip install -e ".[cpu]" --extra-index-url https://download.pytorch.org/whl/cpu
The model's tokenizer needs transformers 5 or newer. Check with
uv run python -c "import transformers; print(transformers.__version__)"
and run `uv pip install -U transformers` if it prints 4.x.
3. Start the server in the background and keep its log:
GPU: SYSONE_MODEL=Okura66/Kahn1-Qwen3.5-4B uv run sysone serve --port 8000
CPU: SYSONE_BACKEND=cpu SYSONE_MODEL=Okura66/Kahn1-Qwen3.5-4B uv run sysone serve --port 8000
For the 3B, set SYSONE_MODEL=Okura66/Kahn1-Qwen2.5-3B instead, on either backend. The engine picks each model's prompt format by itself.
On Windows PowerShell, set each variable first with $env:NAME = "value". Poll GET http://127.0.0.1:8000/health until it returns {"status": "ok"}.
4. POST this JSON to http://127.0.0.1:8000/v1/evaluate/jev (write it to a file and send it with curl -d @file, or use Python). The first call loads the model and can take minutes:
{"state": "Hi, I was charged twice for order #48213. I want a refund, otherwise I will dispute the charge with my bank.",
"schema": {
"intent": {"type": "choice", "instructions": "What is the customer's main request?",
"criteria": {"refund": "the customer wants their money back", "delivery": "question about a delivery", "account": "cannot access their account"}},
"urgency": {"type": "score", "instructions": "How urgent is this message?", "criteria": ["low", "medium", "high", "critical"]},
"churn_risk": {"type": "noul", "instructions": "The customer threatens to leave the service or to escalate."}},
"n_permutations": 1}
Show me the JSON answer and check it: intent has a choice and probabilities that sum to 1, urgency has a level and a score, churn_risk is a number between 0 and 1. Report latency_ms, and send the same request a second time to show the warm latency.
5. Finish with a short summary: backend used, install path, how to stop and restart the server, and where to go next (https://kahn1.com/get-started/ for the API, the response format, tuning and calibration on my own data).
Every link fills the prompt and waits: nothing runs until you read it
and press Enter. If a link does nothing, the app is not installed (Claude Code registers its links after its first
session). From a bash or zsh terminal: claude "$(curl -s https://kahn1.com/get-started/setup-prompt.md)". The prompt is also at
https://kahn1.com/get-started/setup-prompt.md.
02 · Install
Python 3.11 or newer and uv. The GPU backend is vLLM, which needs Linux or WSL2 with CUDA; the CPU backend runs anywhere.
$ git clone https://github.com/Okura66/kahn1 && cd kahn1 $ uv venv --python 3.11 $ source .venv/bin/activate # .venv\Scripts\activate on Windows $ uv pip install -e ".[gpu]" # vLLM, CUDA GPU $ uv pip install -e ".[cpu]" --extra-index-url https://download.pytorch.org/whl/cpu # or CPU only
Weights come from Hugging Face on first use, as a merged model (what vLLM serves) or a LoRA adapter. Kahn1 comes in two sizes, and both run on the GPU and the CPU backends. Kahn1 4B, much stronger on hard judgment calls: the merged model or the LoRA adapter (57 MB, on top of Qwen/Qwen3.5-4B); its merged model is 8.4 GB and needs 12 GB of VRAM or more on vLLM. Kahn1 3B, smaller and faster (median 36 ms against 89 ms): the merged model (6.17 GB) or the LoRA adapter (239 MB, on top of Qwen2.5-3B-Instruct). Code MIT; Kahn1 4B weights Apache 2.0; Kahn1 3B weights under the Qwen Research License of its base model, a research licence: see its terms.
03 · Serve
$ SYSONE_MODEL=Okura66/Kahn1-Qwen3.5-4B uv run sysone serve --port 8000 $ SYSONE_BACKEND=cpu SYSONE_MODEL=Okura66/Kahn1-Qwen3.5-4B uv run sysone serve --port 8000 # no GPU $ curl http://127.0.0.1:8000/health
The model loads on the first request (set SYSONE_EAGER=1 to load it at startup). The server
listens on 127.0.0.1 by default; --host 0.0.0.0 exposes it. Okura66/Kahn1-Qwen2.5-3B
works the same way on either backend. The engine picks the prompt format
from the model (EngineConfig.prompt_format defaults to "auto"): the native chat
template for the Qwen3.5 4B, the tag layout for the 3B.
| Variable | Default | Meaning |
|---|---|---|
SYSONE_MODEL | checkpoints/qwen_merged if present | Hugging Face id or local path |
SYSONE_BACKEND | vllm | vllm (GPU) or cpu (transformers) |
SYSONE_CALIBRATION | calibration.json | temperature file, applied when it exists |
SYSONE_DTYPE | float32 | CPU only; keep float32, bfloat16 is ~7× slower without AMX |
SYSONE_THREADS | all cores | CPU only: torch threads |
SYSONE_EAGER | off | 1 loads the model at startup |
SYSONE_CORS_ORIGIN_REGEX | localhost, kahn1.com | browser origins allowed to call the API |
04 · Send a request
POST /v1/evaluate takes one state and a list of questions. Each question has a
kind and a key that names its answer.
{
"state": "Hi, I was charged twice for order #48213 placed on September 12. I have been waiting a week for customer service to reply. I want a refund right away, otherwise I will dispute the charge with my bank and close my account.",
"questions": [
{
"kind": "choice",
"key": "intent",
"prompt": "What is the customer's main request?",
"options": ["refund: the customer wants their money back", "delivery: question about a delivery", "account: cannot access their account", "information: product information request"],
"allow_other": true
},
{
"kind": "score",
"key": "urgency",
"prompt": "How urgent is this message?",
"levels": ["low", "medium", "high", "critical"]
},
{
"kind": "noul",
"key": "churn_risk",
"statement": "The customer threatens to leave the service or to escalate."
}
],
"n_permutations": 1
}
$ curl -X POST http://127.0.0.1:8000/v1/evaluate \
-H "Content-Type: application/json" -d @request.json
Already have JEV / TypeSafe schemas? POST /v1/evaluate/jev takes the same dictionary
(type, instructions, criteria). /v1/evaluate also
accepts it in a schema field.
$ curl -X POST http://127.0.0.1:8000/v1/evaluate/jev \
-H "Content-Type: application/json" \
-d '{
"state": "I cannot log in to my account since this morning.",
"schema": {
"category": {"type": "choice", "instructions": "Support ticket category",
"criteria": {"bug": "Something is broken", "billing": "Invoices, refunds", "account": "Login, permissions"}},
"urgency": {"type": "score", "instructions": "Level of urgency", "criteria": ["low", "medium", "critical"]},
"legal_threat": {"type": "noul", "instructions": "The customer makes a legal threat"}
}
}'
05 · Read the response
The answer to the request above, recorded from the published 3B on a CPU (hence the latency: on a GPU with vLLM, expect a median of ~36 ms for the 3B and ~89 ms for the 4B at k = 3).
{
"answers": {
"intent": {
"choice": "refund: the customer wants their money back",
"probabilities": {
"refund: the customer wants their money back": 0.9561,
"delivery: question about a delivery": 0.0024,
"account: cannot access their account": 0.0003,
"information: product information request": 0.0002,
"None of these answers": 0.041
},
"confidence": 0.9452
},
"urgency": {
"score": 2.2551,
"level": "high",
"probabilities": {
"low": 0.0013,
"medium": 0.0028,
"high": 0.7356,
"critical": 0.2604
},
"confidence": 0.6474,
"expected_score": 3.2551
},
"churn_risk": {
"noul": 0.9843
}
},
"latency_ms": 21575.8,
"cache_hit_rate": 0.0
}
| Field | Meaning |
|---|---|
choice | the most probable option, as written in the request |
probabilities | the full distribution, sums to 1. Choice adds None of these answers when allow_other is true |
confidence | (p_max − 1/n) / (1 − 1/n): distance from a uniform guess, not the probability of being right |
score / expected_score | Σ i·pi, 0-based / 1-based: a continuous position on the scale |
level | the most probable level |
noul | p(yes), between 0 and 1 |
cache_hit_rate | share of prompt tokens served from the prefix cache (vLLM) |
06 · From Python
from sysone.calibrate import CalibratedEngine, TemperatureConfig
from sysone.engine import Engine, EngineConfig
from sysone.types import ChoiceQuestion, NoulQuestion, Query, ScoreQuestion
engine = Engine(EngineConfig(model="Okura66/Kahn1-Qwen3.5-4B")) # vLLM, GPU
# from sysone.cpu import CPUEngine
# engine = CPUEngine("Okura66/Kahn1-Qwen3.5-4B", dtype="float32") # or a CPU
# "Okura66/Kahn1-Qwen2.5-3B" works the same way in both
engine = CalibratedEngine(engine, TemperatureConfig.load("calibration.json"))
query = Query(state="Hi, I was charged twice for order #48213 ...", questions=[
ChoiceQuestion(key="intent", prompt="What is the customer's main request?",
options=["refund", "delivery", "account", "information"]),
ScoreQuestion(key="urgency", prompt="How urgent is this message?",
levels=["low", "medium", "high", "critical"]),
NoulQuestion(key="churn_risk", statement="The customer threatens to leave."),
])
# Or from a JEV dictionary: Query.from_jev(state=..., schema={...})
resp = engine.evaluate(query, n_permutations=3)
print(resp.answers["intent"].choice, resp.answers["intent"].probabilities)
print(resp.answers["urgency"].expected_score, resp.answers["churn_risk"].noul)
Engine.evaluate_batch(queries) runs several states in one call (the Snake benchmark
advances 20 games that way).
07 · Tune the answer
n_permutations(k, default 3). Choice is asked k times with the options shuffled and the distributions averaged, which cancels position bias. Score, from k = 2, is asked in ascending and descending order. Noul is asked once. Each permutation is one more prompt: on a short state k = 3 costs close to 3×, less on a long state whose prefix the engine caches; k = 1 is the fastest.allow_other(Choice, default true) adds None of these answers as a last option. Set it to false when one option always applies.- More than 26 options. The direct path reads letters A to Z. Above that, call
engine.evaluate_two_stage(query)from Python: one Noul per option shortlists 10 candidates, then a Choice picks among them. The HTTP API does not route to it. - Write options as short labels with a description (
"refund: the customer wants their money back"): the model reads the text, and JEV criteria dictionaries are turned into that form.
08 · Calibrate on your data
Temperatures belong to one model, and Kahn1's were fitted on a split of the training distribution. On your domain, fit your own for the model you serve before you trust a threshold. You need labelled examples in JSONL:
# one labelled example per line; label = index of the right option or level,
# and for noul 1 = yes, 0 = no
{"state": "...", "kind": "choice", "prompt": "...", "options": ["a", "b", "c"], "label": 2}
{"state": "...", "kind": "score", "prompt": "...", "levels": ["low", "mid", "high"], "label": 0}
{"state": "...", "kind": "noul", "statement": "...", "label": 1}
$ uv run sysone calibrate --data my_examples.jsonl --out calibration.json --model Okura66/Kahn1-Qwen3.5-4B $ curl -X POST "http://127.0.0.1:8000/v1/calibrate/load?path=calibration.json" # hot reload
Then pick your automation threshold on held-out examples of your own, and send the rest to a person or a larger model.
09 · In the browser
The playground and Snake run a 4-bit GGUF of Kahn1 3B (mradermacher/Kahn1-Qwen2.5-3B-GGUF, 1.6 to 1.9 GB) with wllama, llama.cpp compiled to WebAssembly and WebGPU. The same prompt and answer code runs in JavaScript. Chrome or Edge with WebGPU: a few seconds per question after the download. Without WebGPU, the CPU fallback takes minutes. Those GGUF quantizations are a smaller build of the 3B, older than the published 3B weights, and not of the 4B.
10 · Go further
To train on your own domain or on another open model (Llama 3.2 3B, Qwen2.5 7B…), follow the
playbook: dataset
format, LoRA training, merging for vLLM, calibration. Evaluation reports and scripts are in
reports/.