Imagine you're a chess player who has to announce your move the instant you see the board. You'd do okay — pattern recognition is fast — but you'd miss the tricky positions. Now imagine you get 30 seconds to think before committing. Your blitz rating goes up, especially on the hard puzzles. That's what Jeeves does to Jev-style decision classifiers: it inserts a reasoning chain between the prompt and the pointer-head decision, and the accuracy gap shows up exactly where you'd expect — on hard and out-of-domain items. The committed claim: a 9B Qwen3.5 model with LoRA and a pointer head, trained with SFT then CISPO (a reinforcement learning schedule), achieves 0.889 on held-out test data versus 0.857 for Jev and 0.822 for Kev-9B. On JevBench hard — the 111 public items designed to stress classifiers — Jeeves hits 0.865 against Jev's 0.730. These are not small margins. The model also produces calibrated probabilities with an ECE of 0.037, better than Jev's 0.049. The architecture is specific and well-documented. Qwen3.5-9B serves as the backbone, augmented with LoRA (rank 16 on all projections) and a pointer head that scores answer options via scaled dot products between query and key projections of hidden states at special tokens. Training follows a two-stage recipe: SFT on 19,126 questions from 12 public datasets plus synthetic policy data, then 624 steps of CISPO (stopped at step 402 for calibration reasons) with 9,992 RL questions and 8 rollouts each. The authors repurpose rare Qwen tokenizer tokens as structural markers — an ablation shows this outperforms plain-text labels. A block-4 diffusion drafter inspired by Orthrus provides 1.6× speedup over greedy decoding, adapted to handle Qwen3.5's Gated DeltaNet layers. The integrity picture is mixed in the way open-source ML work often is. The full training code, data prep scripts, and weights are released. Evaluation uses a mix of public benchmarks (MMLU, SciQ, QNLI, PAWS, TweetEval, Emotion) plus JevBench's public tiers and held-out synthetic rule structures. However, the Kev and Jev comparison numbers come from Kev's own published results on different items from the same sources — not a controlled head-to-head on identical inputs. The JevBench sealed judge tier is excluded. And Jeeves trails Jev substantially on knowledge-heavy benchmarks: MMLU 0.793 vs 0.900, MMLU-Pro 0.739 vs 0.840. The authors are honest about this, but it means the headline accuracy number is doing work that the per-benchmark breakdown complicates. The real question this work participates in: should classification pipelines use reasoning models as an expensive fallback, or can you bake reasoning directly into the classifier? Jeeves argues for the latter. The latency cost is real — 3.3s median with thinking versus 0.3s without, and 17s at p90 for full chains — but the configurable maxthink and nothinkthreshold knobs let you trade accuracy for speed in production. The no-think checkpoint still scores 0.804 on the test split, which sits between Kev-9B (0.822) and Jev (0.857). The obvious experiment not run: scaling to a larger base model. The Qwen3.5 family goes well beyond 9B, and the knowledge-benchmark gap (MMLU, MMLU-Pro) strongly suggests that a larger backbone would close it. The authors likely stopped at 9B because 8×H100 training is already expensive, and doubling parameters would require either more GPUs or a fundamentally different parallelism strategy. The drafter architecture would also need revalidation. This is almost certainly a compute-budget constraint rather than a strategic omission. For practitioners running Jev-style classification pipelines today, Jeeves is immediately useful: it's API-compatible, ships with weights and a Python SDK, and the accuracy improvement on hard items is substantial enough to matter in production routing and escalation decisions. The calibration improvement (ECE 0.037) is arguably more important than raw accuracy for downstream decision-making. The limitation to CUDA Hopper GPUs for the FP8 kernel narrows the deployment surface, but that's the hardware most inference teams are targeting anyway.