Imagine you're at a restaurant with a 200-page menu, but you already know you want one of five dishes. Instead of reading every page, you cover everything except those five entries and point at the one that looks best. That's constrained decoding: mask the entire vocabulary except your allowed answers, run one forward pass, and read off the probabilities. The author demonstrates this with Qwen3-1.7B on multiple-choice questions, turning a generative model into a single-pass classifier. The core mechanism is vocabulary masking at the logit level. After the model processes the prompt in a single prefill pass, you extract the final-position logits, zero out everything except the token IDs corresponding to your answer options (A through E), apply softmax over the survivors, and pick the argmax. One pass, one answer. The technique itself is well-known in the constrained decoding literature — what this post contributes is a clean, reproducible walkthrough with real evaluation numbers. On a random holdout of CommonsenseQA, the base Qwen3-1.7B model scores 59.4% accuracy (725/1221) with a macro F1 of 0.584. A quick fine-tune on the same dataset bumps that to 62.4% accuracy and 0.623 macro F1. These are honest numbers for a 1.7B parameter model on a genuine commonsense reasoning benchmark — not cherry-picked toy examples. The most valuable section is the calibration analysis. The author bins predictions by confidence and reveals that the base model is systematically overconfident: when it predicts with 90-100% confidence, it's only correct 70% of the time. In the 80-90% bin, accuracy drops to ~47%. This is the classic finding from Guo et al. (2017) reproduced in miniature — raw softmax probabilities from neural networks are not calibrated probabilities. The fix applied is temperature scaling, a post-hoc calibration technique. By fitting a single scalar temperature parameter (found to be T=3.797) to minimize the gap between confidence and accuracy, the author achieves dramatically better calibration: the 90-100% confidence bin now corresponds to 95.4% accuracy, and the 70-80% bin hits 77.1%. Temperature scaling is the simplest member of the Platt scaling family and works because it doesn't change the ranking of predictions — only the distribution spread. What this post is NOT is a research advance. Constrained decoding, vocabulary masking, and temperature scaling are all established techniques. The contribution is pedagogical: a clean end-to-end pipeline (dataset → eval → finetune → calibrate) packaged in a GitHub repo with runnable scripts. For practitioners building classification or routing systems on top of LLMs, this is a useful reference implementation. The honest limitation: 59-62% on CommonsenseQA is below current state-of-the-art for models of any serious size, and the calibration technique — while effective — is the simplest possible approach. Isotonic regression or Bayesian binning would likely do better. The post doesn't compare against these alternatives, and it doesn't test on models larger than 1.7B where the dynamics may differ substantially.