Kev
Jared Palmer's open Jev-like Qwen decision models: System One-compatible serve, HF weights 0.8B to 27B, eval tables vs hosted Jev.
Overview
Kev (github.com/jaredpalmer/kev, Apache-2.0) is Jared Palmer's family of small decision models on Qwen3.5 and Qwen3.8 bases (0.8B, 4B, 9B, 27B). Weights and frozen eval suites live on Hugging Face; kev.serve exposes POST /v1/systemone matching TypeSafe's System One contract so the TypeSafe Python SDK can point at localhost. README publishes dev and test accuracy and Brier scores against hosted Jev on new sources (datasets Kev never trained on) and trained sources (held-out rows from Kev training). Kev-27B is within about one point of Jev on new-source accuracy in those tables; Kev-4B is the default starting size. Modal skills, one-command HTTPS deploy, and HF Spaces demo round out the train-and-serve story.
Problem: Teams want Jev-shaped Choice, Score, and Noul gates without per-call TypeSafe bills, but also without brittle chat JSON from general LLMs.
Built for: Engineers who can host Python 3.12+, want System One-compatible APIs, and may fine-tune Qwen3.5 or Qwen3.8 checkpoints on their own labels.
First indexed on Jev Directory: 2026-09-24
Creator and team
- Name
- Jared Palmer
- Handle
- @jaredpalmer
How Jev is used
- Role in the product flow
- Self-hosted System One parallel questions over shared state with calibrated probabilities
- Primitives
- ChoiceScoreNoul
- State in
- Single state string plus a questions map of choice, noul, or score fields per README curl and typesafe_sdk examples.
- Decision out
- Per-question typed answers with probabilities or scores; latency_ms and token usage in the JSON envelope.
- uv sync --extra serve and python -m kev.serve --run jaredpalmer/kev-4b
- POST /v1/systemone with model kev-latest and parallel questions
- Route high-confidence departments or escalate low-confidence tuples in app code
- Optional: fine-tune adapters and redeploy with the repo training and Modal skill paths
Kev is the open-weights answer when you want the same SDK calls as api.typesafe.ai but the bill and GPUs are yours. PostHog Jeeves (posthog-jeeves) adds an optional reasoning chain before the same Choice, Noul, and Score head, which the Jeeves README says helps on JevBench hard but costs latency at the tail. togethercomputer-tev1 teaches a letter-picker via LoRA on Together; featherless-simple-jev assembles logits from arbitrary HF models; bespokelabsai-nimble ships a curated 9B recipe. Kev ships full checkpoints with eval tables against hosted Jev and explicit warnings that Jev training data is unknown, so headline accuracy is directional. wfzyx-von chases sub-25 ms non-autoregressive inference on a smaller encoder; Kev stays autoregressive Qwen with richer Score and Noul in one batch. Do not confuse this repo with kevthetech143-super-jev fan forks. classifier.dev and hosted System One remain the paths when you refuse to operate inference.
Sourced performance claims
- README table lists Kev-4B new-source development accuracy 0.817 and test 0.838 versus Jev hosted 0.857 on development new sources only.Source: github.com/jaredpalmer/kev README Models section
- Public GitHub repo jaredpalmer/kev had 6740 stars and 389 forks when this listing was drafted.Source: GitHub API September 2026
- Apache-2.0 license; default serve example reports about 495 ms latency for a three-question ticket on Kev-4B bf16 on Apple M5 in README.Source: github.com/jaredpalmer/kev README Quick Start
Features and stack
Features
- Four public sizes from 0.8B laptop class to 27B datacenter GPU
- Drop-in TypeSafe Python SDK against local kev.serve
- Frozen HF eval suites and per-model cards with Brier scores
- HF Space demo, GitHub release checksums, Modal fine-tune skill
- Parallel choice, noul, and score questions on one state string
Stack
- Python 3.12 or 3.13
- uv
- Qwen3.5 and Qwen3.8 bases
- CUDA, ROCm, or MLX serve paths
- typesafe_sdk
Pricing: Open source weights; you pay for GPUs, Modal training, or your own cloud serve. No TypeSafe meter unless you call both.
Links
FAQ
- Is Kev the same product as TypeSafe Jev?
- No. Kev reimplements a Jev-like decision stack on open Qwen weights with a compatible HTTP API. Hosted Jev stays on TypeSafe unless you point SDK clients at your server.
- Which size should I start with?
- README recommends Kev-4B for most GPUs, Kev-0.8B when size matters, and Kev-27B when you have 80 GB VRAM and want the best Kev accuracy in their tables.
- How does Kev compare to tev1 or Nimble?
- tev1 is a Together fine-tune recipe; Nimble is Bespoke's 9B schema classifier. Kev ships multiple finished checkpoints with System One parity and Jared Palmer's eval narrative against hosted Jev.
- How is Kev different from PostHog Jeeves?
- Kev answers from the prompt in one forward pass. Jeeves can write a reasoning chain first, then score options. The Jeeves README credits Kev as inspiration and publishes side by side tables where Jeeves leads on some JevBench tiers but trails Jev on MMLU.
Related learn guides
Original Jev guidance that pairs with this product pattern.
- System One model
The model family behind Jev: parallel typed questions, one forward pass, probabilities you can threshold.
- Jev vs LLM classification
When to gate with System One probabilities instead of asking a chat model to label things.
Related products
Hand-picked neighbors with rich profiles or overlapping tags.