Jev.aitools.fyi

Command Palette

Search for a command to run...

Product profile
Python414 starsUpdated 2026-10-07

Jeeves

PostHog's open 9B Jev-like model that reasons before it scores options, with CISPO training, a block-4 drafter, and Jev-compatible POST /v1/systemone.

Overview

Jeeves (github.com/PostHog/jeeves, MIT) is PostHog's open reasoning decision model built on Qwen3.5-9B with LoRA r16 and a pointer head. Nicholas P. Waltz is named as author in the README citation block. The model writes a reasoning chain per question, then scores Choice, Noul, and Score in one POST /v1/systemone call. Training is SFT plus CISPO RL with one fitted temperature and a block-4 diffusion drafter for speed. Weights ship on Hugging Face as PostHog/jeeves (bf16) and PostHog/jeeves-fp8; the repo also publishes jeeves_sdk as a drop-in for the TypeSafe Python SDK. The README credits jaredpalmer/kev as inspiration.

On tables the project publishes with thinking enabled and a 2,560-token cap, Jeeves reports test overall 0.889 versus Jev 0.857 and Kev-9B 0.822, and JevBench overall 0.935 versus Jev 0.866 on 231 public items. The same README shows Jeeves trailing Jev on MMLU (0.793 vs 0.900) and MMLU-Pro (0.739 vs 0.840). Treat every number as the project's own claim, not this directory's measurement.

Problem: Jev-like models score options fast, but accuracy on hard out-of-domain tasks still pushes teams toward slow reasoning LLM fallbacks.

Built for: ML engineers who want an open 9B decision server with optional thinking chains, Jev-shaped HTTP, and published training code on CUDA or Apple Silicon.

First indexed on Jev Directory: 2026-09-29

Creator and team

Name
PostHog
Organization
PostHog
Handle
@PostHog
“Inspired by Kev.”
PostHog/jeeves README Acknowledgements, source

How Jev is used

Role in the product flow
Optional reasoning chain, then parallel Choice, Noul, and Score over shared state via pointer-head softmax
Primitives
ChoiceNoulScore
State in
State string and questions map in TypeSafe System One shape; optional options.think, max_think, nothink_threshold, return_reasoning.
Decision out
Typed answers with probabilities or scores; usage includes reasoning_tokens when thinking is on.
  1. pip install -r requirements.txt and hf download PostHog/jeeves
  2. python -m inference.serve with optional --precision fp8 and drafter_k4 weights
  3. POST /v1/systemone with parallel questions; set options.max_think to cap chain length
  4. Point jeeves_sdk TypeSafeClient at JEEVES_BASE_URL or port 8009 default

Jeeves is the open model that reasons before it decides. jaredpalmer-kev answers straight from the prompt with Qwen fine-tunes and SDK parity tables. liuziyu77-valen adds vision tensors. theoleecj-semif and ollaya-dev-ollaya cover other local serve paths. Hosted TypeSafe Jev remains the closed API baseline in third-party tables. Jeeves serves the same POST /v1/systemone envelope with extra options Jev clients may ignore. Link to Kev both ways when you compare training cost versus chain latency.

Sourced performance claims

  • README results table with thinking: test overall 0.889 vs Jev 0.857 and Kev-9B 0.822; JevBench overall 0.935 vs Jev 0.866 on 231 public items; JevBench hard 0.865 vs Jev 0.730. Kev-9B JevBench cells are Kev-8B per README footnote.Source: github.com/PostHog/jeeves README Results
  • README limitations: MMLU 0.793 vs Jev 0.900; MMLU-Pro 0.739 vs Jev 0.840; thinking median 3.3 s and p90 17.1 s on one H100 fp8 vs about 0.3 s without thinking on 325 dev questions.Source: github.com/PostHog/jeeves README Limitations and Options tables
  • GitHub repo PostHog/jeeves had 414 stars and 21 forks on 2026-10-07. Repo LICENSE is MIT. Hugging Face PostHog/jeeves and PostHog/jeeves-fp8 card license is Apache-2.0.Source: GitHub and Hugging Face API, 2026-10-07

Features and stack

Features

  • Full SFT, CISPO, calibration, and drafter training scripts in-repo
  • Block-4 diffusion drafter with published tokens per second table
  • jeeves_sdk drop-in for typesafe-sdk with reasoning options
  • CUDA bf16 or fp8 and Apple Silicon MPS paths documented
  • export.py and export_fp8.py for fused standalone checkpoints

Stack

  • Python 3.12
  • PyTorch
  • Qwen3.5-9B
  • CUDA 8.9+ for fp8 or Apple Silicon 48 GB+ for bf16 serve
  • Hugging Face weights

Pricing: Open weights and MIT code; you pay for GPUs and electricity. Hugging Face weight license is Apache-2.0 per model cards.

FAQ

Does the Hugging Face license match the GitHub license?
The GitHub repository LICENSE file is MIT. Hugging Face lists Apache-2.0 on PostHog/jeeves and PostHog/jeeves-fp8. Read both before you ship weights.
How do I turn thinking off for latency?
Send options.think false or use nothink_threshold. The README reports about 0.3 s per request without thinking on dev questions versus 3.3 s median with full chains on one H100 fp8.
Are the Kev and Jev comparison tables apples to apples?
The README says comparisons outside JevBench use different items from the same sources. JevBench rows are restricted to the same 231 public items. Kev-9B JevBench numbers are actually Kev-8B.

Related learn guides

Original Jev guidance that pairs with this product pattern.

  • System One model

    The model family behind Jev: parallel typed questions, one forward pass, probabilities you can threshold.

  • Jev vs LLM classification

    When to gate with System One probabilities instead of asking a chat model to label things.

Related products

Hand-picked neighbors with rich profiles or overlapping tags.

IntegrationsPython6,740
Demo
Jared Palmer's open Jev-like Qwen decision models: System One-compatible serve, HF weights 0.8B to 27B, eval tables vs hosted Jev.
integrationsopen-sourceself-hosted
BenchmarksPython24
Reproduce the public Decision Index 0.3 suite locally or as one Hugging Face Job: 37 chance-corrected benchmarks, HTTP /v1/systemone engines, resumable runs.
benchmarksopen-sourceleaderboard
BenchmarksPython3,844
DemoFeatured
Semantic ifs on a home 3090: open models doing structured branches without pretending to be TypeSafe.
benchmarksopen-sourceself-hosted
Agent toolingRust153
Ollama for decision models: pull laya kev von decider, serve POST /v1/systemone on :11435, ONNX graphs only, TypeSafe SDK drop-in.
agent-toolingself-hostedonnx
By Rishit