Start with Kev-4B. Use Kev-9B when accuracy and calibration matter more than memory. Use Kev-0.8B if you need the smallest model. All three are built on Qwen3.5 bases with the same training data and settings.
Each cell is development / test. “Trained sources” means held-out examples from the datasets used to train Kev. “New sources” means datasets and policy rule types Kev wasn’t trained on. Every model was evaluated on the same development sets (decision-v7, transfer-v4) and the same test sets, which were read once per released checkpoint, after model selection. Lower Brier is better.
Kev-9B trails Jev by 3.5 points on the new-source development set (0.822 vs 0.857) and scores 0.852 on the test set, which Jev hasn’t been run on. We don’t know which datasets Jev was trained on, so this isn’t a controlled comparison of the two architectures.
All three models were Two optional settings change the probabilities without changing any answer:
All weights are in the Kev collection and the GitHub release, which includes tarballs and SHA-256 checksums.
The first Kev family used Qwen3 bases with the same data and settings. Those weights stay Because only the base changed, the two generations are a controlled comparison. On the development set the accuracy gain is within noise; on the test set Kev-9B is 7.3 points ahead of Kev-8B (95% CI +2.8 to +11.7) with a Brier score 0.08 lower, Kev-4B is 2.9 points ahead of its predecessor (−0.9 to +6.4), and Kev-0.8B is 4.8 points ahead of Kev-0.6B (+0.2 to +9.3). PLAN_Qwen35.md has the full experiment, including the criteria we set in advance and how the results measured against them.
The original Kev-0.5B used Qwen2.5-0.5B and is kept for reference; see its model card.
POST /v1/systemone
state is the text to evaluate. Each question has instructions and, where needed, a set of answers to choose from.
{
“state”: “…”, // string | object | array — the content to evaluate
“model”: “kev-latest”,
“questions”: {
“
Type Criteria Answer
noul
Optional descriptions for true and false
noul: probability of yes
choice
1–255 option names, each with a description or null
choice: most likely option; probabilities and confidence
score
2–255 descriptions, ordered from lowest to highest
score: mean level index, starting at 0; legend, probabilities, and confidence
For Choice with K > 1 options, confidence is (p_max − 1/K) / (1 − 1/K). A single option has confidence 1. Score confidence measures how close the distribution is to its most likely level. It’s an approximation of TypeSafe’s formula, which isn’t public. Neither field is a measured accuracy rate.
Objects and arrays are converted to labeled text. Delimiter-like strings in user input are escaped before tokenization. Invalid requests return 422. usage.output_tokens counts tokens in the serialized answers, not generated tokens.
The server binds to 127.0.0.1 and has no authentication. Keep it local unless you add authentication yourself.
Each checkpoint is a rank-16 LoRA adapter and a small pointer head on a Qwen base model. On an attention-only base (Qwen3), the state and questions go into one token sequence:
<state> …state… <q> instructions <opt> option 1 </opt> <opt> option 2 </opt> … <decide> <q> instructions <opt> option 1 </opt> <opt> option 2 </opt> … <decide>
The attention mask lets a token read the state and its own question, but not other questions or future tokens. Each question’s position IDs restart just after the state. This lets the model process the state once and answer each question independently.
Qwen3.5 mixes attention layers with Gated DeltaNet layers, which are recurrent and ignore attention masks. For those models, each question runs as its own row: the state followed by that question, with the same positions as above. The rows are independent, so isolation is exact, and the server computes the state once and reuses its cache for every row. On attention-only models the two forms give identical probabilities (tests/test_v3.py).
The pointer head scores each option’s </opt> hidden state against the question’s <decide> hidden state. A softmax turns those scores into probabilities. Because <decide> comes last, it can attend to the full option list.
Training uses cross-entropy on the correct answer. The adapter and head are trained together; the rest of the base weights stay fixed. Training examples and API requests use the same text format. No Jev outputs were used for training.
Asking questions together or separately produces probabilities within 4e-6 in the fp32 tests. This does not mean option order is irrelevant: options within a question can still affect one another. See the model code and parity tests.
On CUDA, install flash-linear-attention for the Qwen3.5 models (the Modal image does this); a five-question request takes tens of milliseconds on an H100.
On Apple Silicon there are no fast kernels for the DeltaNet layers, so PyTorch runs reference code. Median model time in bf16 on an M5, five questions with three options each on a ~230-token state:
If you serve on a Mac and need low latency, use the Qwen3 models for now. An MLX backend for the Qwen3.5 models is the next planned change.
For the attention-only models the server merges the LoRA weights in fp32 before casting, uses SDPA attention on Apple GPUs, pads MPS inputs to 64-token buckets, and caches the state prefix for repeated requests (four states of at least 384 tokens by default). With a repeated 772-token state, Kev-4B (Qwen3) answers in 242 ms instead of 861 ms.
You can disable these with KEV_MERGE=0, KEV_ATTN=eager, KEV_SHAPE_BUCKET=1, and KEV_PREFIX_CACHE=0. On 24 new-source records, bf16 probabilities differed from fp32 by at most 0.017, with no change in the highest-probability answer. That is a small check, not a guarantee for every input.
The released models use decision-v7: 10,000 examples from ten public datasets, 896 generated policy examples, and 1,680 examples from 60 generated rule structures. All train for two epochs with LoRA rank 16 and cross-entropy. The learning rate is 1e-4 for 0.8B and 5e-5 for 4B/9B. For Qwen3.5 bases the adapter also covers the DeltaNet projections; kev.train picks the right targets from the model config.
sanity run, ~1 minute
uv run python -m kev.train –n_per_source 40 –accum 4 –out runs/smoke
Kev-0.8B (~20 min on one H100; the Mac path works but is slow for Qwen3.5 bases)
uv run python -m kev.train –suite evals/v7/decision-v7 –base Qwen/Qwen3.5-0.8B-Base –base_revision dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68
–epochs 2 –lr 1e-4 –batch 8 –dtype bf16 –p_none_pair 0.25 –device cuda –out runs/kev-0.8b
the Kev-4B recipe (one H100 via Modal, ~1 h; see below). Swap in Qwen/Qwen3-4B-Base for the previous generation.
uv run python -m kev.train –suite evals/v7/decision-v7 –base Qwen/Qwen3.5-4B-Base –base_revision 1001bb4d826a52d1f399e183466143f4da7b741b
–epochs 2 –lr 5e-5 –batch 4 –accum 2 –dtype bf16 –checkpointing 1 –p_none_pair 0.25 –device cuda –out runs/kev-4b
Fine-tuning on your own data
The released models were trained on public datasets and generated policy examples. If your questions look different — your own routing categories, your own escalation rules, another language — a short fine-tune on a few hundred labelled examples usually helps more than any prompt change.
Put your examples in a JSONL file, one request per line. It’s the same shape as an API request, plus a label on every question:
{“state”: {“subject”: “Charged twice”, “body”: “I see two charges for order #4411. Please refund one.”}, “questions”: { “team”: {“type”: “choice”, “instructions”: “Which team should handle this ticket?”, “criteria”: {“billing”: “Payments and refunds”, “shipping”: “Delivery problems”, “access”: “Login and account access”}, “label”: “billing”}, “angry”: {“type”: “noul”, “instructions”: “Is the customer angry?”, “label”: false}, “priority”: {“type”: “score”, “instructions”: “How urgent is this ticket?”, “criteria”: [“low”, “normal”, “high”], “label”: 1}}}
For choice the label is the option name, for noul it’s true or false, and for score it’s the level’s position starting at 0. Keep 10–20% of the file aside for evaluation.
Then start from a released checkpoint with --init_from:
uv run python -m kev.train –data train.jsonl –base Qwen/Qwen3.5-4B-Base –init_from jaredpalmer/kev-4b
–epochs 2 –lr 2e-5 –batch 1 –accum 8 –dtype bf16 –checkpointing 1 –device cuda –out runs/mine
uv run python -m kev.benchmark –run runs/mine –data heldout.jsonl –out runs/mine-eval KEV_DTYPE=bf16 uv run –extra serve python -m kev.serve –run runs/mine –port 8009
--init_from loads the adapter and pointer head from the released model before training, so you keep what Kev already knows and add your domain on top. Starting from the base model instead throws that away: in one user’s test on 836 support-tool decisions, a fine-tune from the base scored 0.33 on Kev’s own evaluation set, against 0.84 for the released model; the same data with --init_from kept 0.83 there and reached 0.88 on the new domain. Use a smaller learning rate than the from-scratch recipe (2e-5 is a good start), and pick --base to match the checkpoint you start from; the trainer checks that the base, revision, LoRA rank, and head size agree before it loads anything.
--batch 1 --accum 8 in bf16 fits the 0.8B model on a 4 GB GPU. The benchmark reports accuracy, Brier score, and calibration per question type, so you can see which of your questions the fine-tune helped. The checkpoint you started from is recorded in runs/mine/training_config.json.
Use uv run python -m kev.train --help for all training options. The released models don’t use the optional --perm_kl or --ord_w losses. The model cards have the training settings and dataset lists; PLAN.md records what was tried and what helped.