PostHog releases Jeeves, an open-source 9B reasoning decision model

PostHog has published Jeeves on GitHub, a 9B model in the style of Jev that thinks before it decides. It is built on Qwen3.5-9B with LoRA and a pointer head, trained with supervised fine-tuning (SFT) followed by CISPO reinforcement learning. The repository ships the full training code, train/dev/test data, a block-4 diffusion drafter for faster inference, and a Python SDK. The README says it was inspired by Kev.
The pitch: Jev-like models give calibrated decision probabilities, but at low accuracy, so many pipelines fall back on a reasoning model. Jeeves trains a Jev-like Qwen3.5-9B to reason first. The README says this gives better performance on out-of-domain tasks and outperforms Jev on JevBench hard (public).
The reported numbers are self-reported. On test data the model never trained on, Jeeves scores 0.889, against 0.822 for Kev-9B and 0.857 for Jev (the table caption reads: accuracy with thinking, greedy, 2,560-token cap). On JevBench's public easy, standard and hard tiers (231 items in total, sealed judge tier excluded) it scores 0.935 against 0.866 for Jev. On its own test split of 2,962 items, the same checkpoint scores 0.804 without thinking and 0.840 with it. The Kev-9B and Jev columns are the numbers Kev publishes. No Kev-9B JevBench result is published, so the JevBench comparison uses Kev-8B (Qwen3).
One request can mix yes/no (noul), multiple-choice (choice) and rating (score) questions, through a Jev-compatible API. Speed is about 0.3 s per request without thinking and a 3.3 s median with it on one H100. Thinking can be sped up by truncating chain length. It runs on CUDA, with Hopper needed for the FP8 kernel.
How it works: the state and questions go into the Qwen chat template, marked with rare, largely unused Qwen tokens ("<|fim_prefix|>", "<|fim_middle|>", "<|box_start|>", "<|box_end|>", "<|fim_suffix|>") rather than plain text. The model writes a reasoning chain, the questions are repeated after the token, and a pointer head scores each option using a scaled dot product between a query projection of the hidden state at
Training had three stages. SFT ran 2 epochs, 596 steps on 8 GPUs, with LoRA r=16 on all projections of Qwen3.5-9B plus the pointer head, on 19,126 questions from 12 public datasets and synthetic policy data; half the questions carry a reasoning chain sampled from the base model. CISPO used a 624-step schedule but was stopped at step 402, on 9,992 RL questions with 8 rollouts each at temperature 1, thinking capped at 2,560 tokens. The README says stopping at 402 keeps the best calibration and dev score, and that past it the head over-sharpens on the saturated RL pool. Calibration is the single fitted temperature, stored with the checkpoint.
The drafter is a diffusion view of the frozen model, inspired by Orthrus. Unlike Orthrus, which supports attention-only models, it supports Qwen3.5's Gated DeltaNet layers by letting mask tokens cross-attend to those layers' post-convolution keys and values. Block 4 is the default because it stays cheap when several questions are batched.
The README lists limits itself. Knowledge questions trail Jev (MMLU 0.793 vs 0.900, MMLU-Pro 0.739 vs 0.840). Thinking is slow at the tail: 17 s at p90 with full chains. Comparisons with Kev and Jev outside JevBench use different items from the same sources. No language consistency reward was used, so thinking chains are not well interpretable. The citation entry names Nicholas P. Waltz as author and gives 2026 as the year.
Key facts
- Jeeves is a 9B Jev-like model on Qwen3.5-9B (LoRA r=16, pointer head), trained with SFT and CISPO to reason before deciding; code, train/dev/test data and a block-4 diffusion drafter are included.
- Self-reported: 0.889 on held-out test data versus 0.822 for Kev-9B and 0.857 for Jev, and 0.935 versus 0.866 for Jev on 231 public JevBench items.
- Thinking helps: 0.840 with thinking versus 0.804 without on the 2,962-item test split; latency is about 0.3 s without and a 3.3 s median with thinking on one H100, but 17 s at p90 with full chains.
- It trails Jev on knowledge benchmarks: MMLU 0.793 vs 0.900 and MMLU-Pro 0.739 vs 0.840.
- A Jev-compatible API and a drop-in Python SDK handle yes/no, multiple-choice and rating questions in one request; it needs Python 3.12 and a CUDA GPU.
Why it matters
The README frames a familiar trade-off: Jev-like models give calibrated decision probabilities but at low accuracy, so many pipelines add a reasoning model as a fallback. Jeeves tries to put reasoning inside the Jev-style classifier itself, keeping probabilistic outputs for yes/no, multiple-choice and rating questions. The reported gain from thinking on its own test split is 0.804 without versus 0.840 with. The release also includes the training recipe, data and a drafter that supports Qwen3.5's Gated DeltaNet layers, which the README says Orthrus does not.
Who it affects
Teams running decision or classification pipelines built on Jev-style models, especially those that already fall back to a separate reasoning model, are the target. The SDK is a drop-in replacement for Jev's Python SDK, and the API is Jev-compatible. Researchers working on CISPO, calibration or diffusion drafters for hybrid Qwen3.5 architectures get the code and data to reproduce the setup.
How to use it
You need Python 3.12 and a CUDA GPU; Hopper is needed for the FP8 kernel. Install the requirements, run hf download PostHog/jeeves --local-dir jeeves-weights, then serve with python -m inference.serve, passing the model, the drafter (drafter_k4.safetensors) and a port. Or fuse your own checkpoint with export.py and serve that. Requests go to /v1/systemone in Jev's format, with a state and a set of named questions of type choice, noul or score. Each answer returns probabilities and a confidence. The optional options field (for example max_think) is ignored by Jev clients that do not send it. The SDK is installed with pip install ./sdk; its client connects to http://127.0.0.1:8009 by default (or JEEVES_BASE_URL), needs no API key and waits up to 120s. When latency matters, the README points to max_think and nothink_threshold. The README's example response on one H100 (FP8), with three questions thinking in parallel, reported 8141.6 ms; that is a single illustration, not a benchmark. The source gives no license for the weights or code; each public dataset stays under its own license.
How solid is it
All results are self-reported by the README, and no independent evaluation or third-party reproduction is mentioned. The README is open about comparison caveats: the Kev-9B and Jev columns are the numbers Kev publishes; the JevBench comparison uses Kev-8B (Qwen3) because no Kev-9B result is published; the sealed judge tier is excluded and the Jev and Kev numbers are restricted to the same public items; and comparisons outside JevBench use different items from the same sources. The 0.889 test figure is labelled accuracy with thinking, greedy, 2,560-token cap; the source does not reconcile it with the 0.840 with-thinking score on the 2,962-item test split. The results tables themselves are not in the source text.
Risks and caveats
Knowledge questions trail Jev: MMLU 0.793 vs 0.900 and MMLU-Pro 0.739 vs 0.840. Thinking is slow at the tail, 17 s at p90 with full chains, against a 3.3 s median. No language consistency reward was used, so the thinking chains are not well interpretable. The CISPO run was stopped at step 402 of 624 because later steps over-sharpened the head on the saturated RL pool, so the result depends on that stopping choice. Hardware is limited to CUDA GPUs, with Hopper for FP8. The source gives no license for the weights or code.
“Jev-like models give calibrated decision probabilities, but at low accuracy.”
— Jeeves README, PostHog/jeeves on GitHub