Kev releases 1.0 open-weight decision models with a 64k document window
Kev 1.0 introduces Kev-27B and updates Kev-9B with a 64k document window and TypeSafe SDK compatibility. The release includes weights, source code, and fine-tuning and deployment skills.

TL;DR
- Kev 1.0 packages four open-weight decision models, including Kev-27B and an updated Kev-9B, with TypeSafe SDK compatibility, according to the launch post.
- Kev-27B is a full fine-tune of Qwen3.8-27B’s text backbone on roughly 146,000 records and 337,000 questions, as detailed in the training announcement.
- Broader-task gains come with a long-contract regression: the follow-up reports improved generalization, while the release write-up describes weaker accuracy and overconfidence on contracts.
- The self-hosting package includes downloadable weights and model cards, plus source code and fine-tuning and deployment skills.
The 27B model card records an 85/15 blend with an older checkpoint and an 87.1 GB memory peak at the advertised context ceiling. The release notes also document a fix for silent document truncation and a kernel change that moved probabilities without changes to Kev’s own code.
Four pinned checkpoints
Jared Palmer, Kev’s developer, released the versioned family on October 1. The release fixes checkpoint, evaluation and serving-code versions; the 27B and 9B updates were already published on September 30.
The release manifest contains:
- Kev-0.8B: Qwen3.5-0.8B-Base, LoRA adapter plus pointer head.
- Kev-4B: Qwen3.5-4B-Base, LoRA adapter plus pointer head.
- Kev-9B v2: Qwen3.5-9B-Base, LoRA adapter plus pointer head.
- Kev-27B v2: Qwen3.8-27B’s post-trained text backbone, full bf16 weights plus pointer head.
All four carry Apache-2.0 licenses and a v1.0 Hub tag. GitHub provides checksum-bearing assets for the three smaller checkpoints; the 51 GB 27B checkpoint lives on Hugging Face because it exceeds GitHub’s per-asset limit.
Typed questions and shared caches
Kev accepts a document, called the state, plus questions and candidate answers. Its System One endpoint, POST /v1/systemone, exposes three primitives:
noul: probability of yes for a yes/no question.choice: probability distribution over named options, the selected option and confidence.score: probability distribution over ordered levels and their expected index.
The server computes the document once and reuses its cache across independent question rows. Exact question isolation is the clever bit: Qwen’s recurrent DeltaNet layers ignore attention masks, so separate rows prevent one question from influencing another.
Palmer’s write-up reports roughly 9.4 seconds to read a 64k-token document on an H200, then roughly 0.7 seconds for subsequent questions using the cache. The Python SDK keeps the same request shape when its endpoint and model name change.
The same API is attracting practical Jev guides across the community.
Full-weight SFT and an 85/15 blend
Kev-27B’s training recipe differs from the smaller models’ frozen-backbone adapters. The published recipe in the model card has three stages:
- Full-weight fine-tuning: one epoch over 145,840 records containing 337,130 questions, with states up to 32,768 tokens, on eight H200 GPUs.
- Weight averaging: 85% of the new fine-tuned backbone plus 15% of the previous checkpoint’s backbone. The new pointer head is retained.
- Calibration: a single temperature,
T = 1.32, fitted on 648 held-out questions.
Training combines public tasks, synthetic decisions and code-generated examples. No Jev outputs were used; the full corpus remains private, with its manifest published.
The “27B” model loads 25.6 billion backbone parameters. Its vision tower, language-model head and multi-token-prediction layers are excluded.
64k documents and an 87 GB peak
The model card defines “validated context” as the longest document length whose contract accuracy remains within three percentage points of the same model at 8k tokens, using a 95% lower confidence bound.
- Kev-27B: validated through 65,536 tokens, despite training states stopping at 32,768.
- Kev-0.8B, 4B and 9B: validated through 8,192 tokens. Their servers accept 65,536, but the contract tests exceed the tolerated accuracy drop beyond 8k.
- Kev-27B memory: about 65.5 GB resident; measured peaks of 78.7 GB at 32k tokens and 87.1 GB at 64k on an H200. The maximum document length needs more than an 80 GB GPU.
The release notes describe a serving fix with direct operational consequences: over-limit states now return HTTP 422 with the token count and limit. Earlier serving code silently cut them; opting back into truncation with KEV_TRUNCATE_STATES=1 now marks responses as truncated.
Policy gains and contract regressions
Palmer’s release write-up reports these Kev-9B changes after one additional fine-tuning pass:
- Policy/rule reasoning: 58% → 83%, +25 points.
- Developer-tool decisions: 64% → 79%, +15 points.
- Consumer-complaint sorting: 83% → 90%, +7 points.
- Held-out dataset index: 40 → 41, +1 index point.
Kev-27B’s model card separates gains on different evaluation panels:
- Audited held-out public datasets: 82.0% → 83.2%, +1.2 points versus v1.
- Skills, tooling and documents: 80.0% → 88.9%, +8.9 points, on held-out items from trained task families.
- Locked out-of-domain test: 89.63% → 88.87%, −0.76 points; the reported interval includes zero.
- CUAD contracts: 89.0% → 87.4%, −1.6 points across tested lengths. Expected calibration error worsened from 0.007 to 0.053.
The older 27B checkpoint remains available as jaredpalmer/kev-27b@v1-lora. Maintaining accuracy across longer inputs and improving accuracy over the previous model are separate measurements here.
Fine-tuning and deployment skills
The fine-tuning skill packages the customization loop for a coding agent:
- Identify questions and collect or generate labeled records.
- Fine-tune from a released checkpoint on Modal.
- Fit a temperature on held-out data and compare against the untouched model.
- Deploy an endpoint and clean up the training resources.
The repository reports a roughly $1 Kev-4B training run on an H100. On its example support workload, 1,050 generated records raised accuracy from 67.7% to 73.6%, and coverage at a 5% error budget from 34% to 48%; a 400-record experiment produced a gain inside the noise.
The deployment skill creates a bearer-authenticated HTTPS endpoint that scales to zero. The first request after idle waits approximately 35 seconds for container startup.
Clef and d1
Cloudflare released Clef and Clef-flash on Workers AI on October 1, with Apache-2.0 weights and Jev-compatible APIs. Clef adds a vision encoder, extending the decision interface to image classification.
Liquid AI announced d1 on September 29. A launch follow-up described the checkpoint as experimental and free to test, with quality, latency and feature improvements still planned.
Kernel provenance and evaluation audits
Kev’s release notes record an evaluation-image kernel change that shifted the previous 27B checkpoint’s probabilities by 0.03–0.06 without changes to Kev’s code. Reports now record package versions, GPU, dtype, attention kernels and DeltaNet implementations.
The release also retired three evaluation sources from model selection:
- Scienthoon: templated tickets, including a question the text could not answer.
- WANLI-v1/v2: roughly a quarter of gold labels came from one of two disagreeing annotators.
- TypeSafe’s public evaluations: labels supplied by two closed models, with too few questions to distinguish checkpoints reliably.
Earlier results remain in the historical records. The notes additionally flag optimistic test margins for Kev-27B v2 and Kev-9B v2 because re-selection happened with earlier development reads already known.