Imagine you're a gymnastics judge. You don't watch all six routines, then rank them against each other — you score each gymnast independently on the same rubric, blind to the others. If the order of performers changed, your scores wouldn't. That's the core mechanism here: ufakzeka-karar scores each answer option at shared token positions, blind to the other options, so shuffling the choices doesn't change the answer. A sequential head trained with shuffled options achieved similar accuracy but still flipped 2.3–2.8% of answers on reorder — a small but measurable leak of positional bias. The committed claim: a 182M-parameter Turkish-language model that takes a question and a set of options (multiple choice, ordinal scale, or yes/no) and returns calibrated probabilities plus an expected-error "not sure" signal — no autoregressive generation, one CPU forward pass, up to ten options. This is not a frontier result in accuracy terms. On HakemBench v1.0 (4,275 questions, 7 tracks), it ranks 7th of 16 with a composite of 0.660. The interesting part is the architecture choice and the integrity disclosure. Architecturally, this is a classification head bolted onto ufakzeka-1-base (the lab's Turkish foundation model), operating in the encoder-as-scorer family rather than the generative-decoder family. Each option is scored at shared positions — meaning the computation is embarrassingly parallel across options and fully order-invariant by construction. Temperature scaling is applied for calibration. The REINFORCE training alternative lost 10.2 points of macro F1 versus cross-entropy, a clean ablation that justifies the training objective choice. The integrity story is where this paper earns real respect. The authors flag that their released model is the third of three runs scored on HakemBench, and the numbers are not blind. The second run's training data was explicitly aimed at the first run's errors on guardrails, moderation, and customer support — meaning test-set signal leaked into training. The released model trained after reading the second run's guardrail results. These tracks carry a written flag: "shaped by reading the test results." When you exclude those contaminated tracks and score only the four clean ones, the composite rises to 0.678, 6th of 16. Temperature scaling also shows an honest limitation: smooth ECE drops on dev but rises on held-out support questions (0.027→0.045), and the released model's calibration numbers (0.036→0.064) aren't unseen-question tests because it trained on them. The ladder context is thin. HakemBench is a new benchmark with 16 entries, but the paper doesn't name or compare against specific Turkish NLU models (e.g., BERTurk, Turkish GPT variants) on shared tasks with shared metrics. The 7th-of-16 ranking is honest but hard to interpret without knowing what the top models are and how much bigger they are. The 182M parameter count is small — this is a lightweight, deployable model, not a SOTA chase. The milestone question is about Turkish-language NLU infrastructure more broadly. At 182M parameters and CPU-only inference, this targets production deployment for structured decision tasks — customer support triage, moderation, survey analysis. The next concrete threshold would be competitive performance on all 7 HakemBench tracks without any test-set contamination, and scaling the option-blind scoring head to larger base models to see if the order-invariance property holds and accuracy climbs. The obvious experiment not run: testing this architecture on a non-Turkish benchmark (e.g., MMLU or a multilingual decision task) to see whether the order-invariant scoring head is a generalizable contribution or a Turkish-specific engineering artifact. The honest read is (a) — the lab is Turkish-focused, the benchmark is Turkish, and cross-lingual experiments are outside their immediate scope and compute budget. The absence doesn't undermine the paper, but it limits the generalizability claim.