Jeeves
PostHog's open 9B Jev-like model that reasons before it scores options, with CISPO training, a block-4 drafter, and Jev-compatible POST /v1/systemone.
Overview
Jeeves (github.com/PostHog/jeeves, MIT) is PostHog's open reasoning decision model built on Qwen3.5-9B with LoRA r16 and a pointer head. Nicholas P. Waltz is named as author in the README citation block. The model writes a reasoning chain per question, then scores Choice, Noul, and Score in one POST /v1/systemone call. Training is SFT plus CISPO RL with one fitted temperature and a block-4 diffusion drafter for speed. Weights ship on Hugging Face as PostHog/jeeves (bf16) and PostHog/jeeves-fp8; the repo also publishes jeeves_sdk as a drop-in for the TypeSafe Python SDK. The README credits jaredpalmer/kev as inspiration.
On tables the project publishes with thinking enabled and a 2,560-token cap, Jeeves reports test overall 0.889 versus Jev 0.857 and Kev-9B 0.822, and JevBench overall 0.935 versus Jev 0.866 on 231 public items. The same README shows Jeeves trailing Jev on MMLU (0.793 vs 0.900) and MMLU-Pro (0.739 vs 0.840). Treat every number as the project's own claim, not this directory's measurement.
Problem: Jev-like models score options fast, but accuracy on hard out-of-domain tasks still pushes teams toward slow reasoning LLM fallbacks.
Built for: ML engineers who want an open 9B decision server with optional thinking chains, Jev-shaped HTTP, and published training code on CUDA or Apple Silicon.
First indexed on Jev Directory: 2026-09-29
Creator and team
- Name
- PostHog
- Organization
- PostHog
- Handle
- @PostHog
“Inspired by Kev.”
How Jev is used
- Role in the product flow
- Optional reasoning chain, then parallel Choice, Noul, and Score over shared state via pointer-head softmax
- Primitives
- ChoiceNoulScore
- State in
- State string and questions map in TypeSafe System One shape; optional options.think, max_think, nothink_threshold, return_reasoning.
- Decision out
- Typed answers with probabilities or scores; usage includes reasoning_tokens when thinking is on.
- pip install -r requirements.txt and hf download PostHog/jeeves
- python -m inference.serve with optional --precision fp8 and drafter_k4 weights
- POST /v1/systemone with parallel questions; set options.max_think to cap chain length
- Point jeeves_sdk TypeSafeClient at JEEVES_BASE_URL or port 8009 default
Jeeves is the open model that reasons before it decides. jaredpalmer-kev answers straight from the prompt with Qwen fine-tunes and SDK parity tables. liuziyu77-valen adds vision tensors. theoleecj-semif and ollaya-dev-ollaya cover other local serve paths. Hosted TypeSafe Jev remains the closed API baseline in third-party tables. Jeeves serves the same POST /v1/systemone envelope with extra options Jev clients may ignore. Link to Kev both ways when you compare training cost versus chain latency.
Sourced performance claims
- README results table with thinking: test overall 0.889 vs Jev 0.857 and Kev-9B 0.822; JevBench overall 0.935 vs Jev 0.866 on 231 public items; JevBench hard 0.865 vs Jev 0.730. Kev-9B JevBench cells are Kev-8B per README footnote.Source: github.com/PostHog/jeeves README Results
- README limitations: MMLU 0.793 vs Jev 0.900; MMLU-Pro 0.739 vs Jev 0.840; thinking median 3.3 s and p90 17.1 s on one H100 fp8 vs about 0.3 s without thinking on 325 dev questions.Source: github.com/PostHog/jeeves README Limitations and Options tables
- GitHub repo PostHog/jeeves had 414 stars and 21 forks on 2026-10-07. Repo LICENSE is MIT. Hugging Face PostHog/jeeves and PostHog/jeeves-fp8 card license is Apache-2.0.Source: GitHub and Hugging Face API, 2026-10-07
Features and stack
Features
- Full SFT, CISPO, calibration, and drafter training scripts in-repo
- Block-4 diffusion drafter with published tokens per second table
- jeeves_sdk drop-in for typesafe-sdk with reasoning options
- CUDA bf16 or fp8 and Apple Silicon MPS paths documented
- export.py and export_fp8.py for fused standalone checkpoints
Stack
- Python 3.12
- PyTorch
- Qwen3.5-9B
- CUDA 8.9+ for fp8 or Apple Silicon 48 GB+ for bf16 serve
- Hugging Face weights
Pricing: Open weights and MIT code; you pay for GPUs and electricity. Hugging Face weight license is Apache-2.0 per model cards.
Links
FAQ
- Does the Hugging Face license match the GitHub license?
- The GitHub repository LICENSE file is MIT. Hugging Face lists Apache-2.0 on PostHog/jeeves and PostHog/jeeves-fp8. Read both before you ship weights.
- How do I turn thinking off for latency?
- Send options.think false or use nothink_threshold. The README reports about 0.3 s per request without thinking on dev questions versus 3.3 s median with full chains on one H100 fp8.
- Are the Kev and Jev comparison tables apples to apples?
- The README says comparisons outside JevBench use different items from the same sources. JevBench rows are restricted to the same 231 public items. Kev-9B JevBench numbers are actually Kev-8B.
Related learn guides
Original Jev guidance that pairs with this product pattern.
- System One model
The model family behind Jev: parallel typed questions, one forward pass, probabilities you can threshold.
- Jev vs LLM classification
When to gate with System One probabilities instead of asking a chat model to label things.
Related products
Hand-picked neighbors with rich profiles or overlapping tags.