Imagine you're at a restaurant with a 200-page menu. A normal waiter reads you the entire menu aloud, improvises descriptions, and eventually gets to what you asked. A decision waiter just points at the three items you narrowed it down to and says 'this one, 77% confident.' Strands Decider 2B is the pointing waiter — it can't compose poetry or summarize documents, but it can pick between options extremely fast and tell you how sure it is. The committed claim: take a pre-trained LLM (Qwen3.5-2B), rip out the language-model head that generates text, replace it with a pointer head (~1M parameters) that scores hidden states at each option position against the hidden state at an position, fine-tune with LoRA rank-16, and you get a model that is faster, more reliable for classification tasks, always produces a valid answer, and provides calibrated confidence scores. This is the second model in a new 'decision model' class that TypeSafe AI's Jev kicked off earlier this month. The ladder context matters here. On JevBench's public set, Strands Decider 2B ranks 3rd of 33 models in the 2B parameter class, and 1st of 30 when you exclude models slightly above the 2B threshold. That's a strong showing for a v19 iteration with full transparency on architecture evolution. But the benchmark itself is young — JevBench is essentially the only game in town for this model class, and there's no independent validation beyond it. The accuracy/calibration Brier score trajectory across 19 versions shows consistent improvement, which is more convincing than a single snapshot. Architecturally, this sits in the encoder-style classification family, not the autoregressive generation family. The pointer head computes a dot-product-like score between option tokens and a special token in the final hidden states — conceptually similar to how extractive QA models pick span boundaries. The LoRA adapter keeps the fine-tuning lightweight. The key hardware property: 2B parameters runs on consumer GPUs (RTX 3090) or Apple Silicon (M3 MacBook), with median latency of 115ms and 153ms respectively. Latency scales approximately linearly with task token count. The integrity picture is unusually strong for an industry lab release. All training data, scripts, code, and model weights are open-sourced. The full version history (v1 through v19) is documented including the failed slot-head architecture. The benchmark is JevBench's public set — community-standard for this nascent model class but still a single benchmark. No pre-registration, no independent replication, and the benchmark was likely known during development. The confidence calibration claim (Brier score) is the most interesting integrity signal: calibration is harder to game than raw accuracy. The practical use cases tell you where this class of model fits: tool selection in agentic workflows, guardrails (should this tool call proceed?), model routing (which LLM handles this query?), sentiment scoring, language detection, policy classification. The worked example in the article — an intervention handler that checks whether a Strands agent's tool-call arguments are grounded in user input before executing — is the pattern to watch. A 115ms decision gate before every tool call is cheap enough to be always-on; an LLM inference call in the same position would be prohibitively slow. The 20-year question for decision models is whether they carve out a permanent architectural niche or get absorbed back into frontier LLMs that learn to do fast classification internally. The efficiency argument is real today: if you can decompose an agentic workflow into many small classification decisions plus a few expensive generation calls, your cost and latency drop dramatically. Whether that decomposition stays valuable as LLMs get cheaper and faster is the open question. For now, Strands Decider 2B is a clean, well-documented contribution that gives developers a concrete tool to experiment with a genuinely new model class.