Imagine you're grading a stack of essays with a 20-item checklist. The naive approach: one point per checked box, sum them up. But you already know some boxes are gimmes ("Is there a title?") and others are brutal filters ("Does the argument survive a steel-man counterexample?"). A student who nails the hard box but misses the easy one is almost certainly better than the reverse — yet the naive sum treats them identically. Item Response Theory (IRT), the psychometric engine behind standardized tests like the GRE, solves exactly this: it models each item's difficulty and discrimination power, then estimates a latent ability score from the response pattern. This paper does IRT for RLHF rubrics. The committed claim: when rubric criteria are monotone indicators of a shared quality axis, a two-parameter IRT model (treating each criterion as a "test item" and each rollout as a "test-taker") produces a reward signal with strictly higher signal-to-noise ratio than summing assigned points. The authors call this Rubric Response Theory (RRT). The key structural insight is that naive point-summing is degenerate — distinct verdict patterns collapse to the same scalar, destroying information the judge already produced. RRT introduces a Response Parameter Network (RPN) that reads prompt and criterion text to predict each criterion's difficulty and discrimination parameters. As the policy improves during training, the item statistics shift — criteria that were hard become easy. Online expectation-maximization updates the RPN from current rollout verdicts, keeping the reward surface calibrated to the moving policy distribution. This is the mechanism that prevents reward hacking from stale difficulty estimates. The ladder comparison is against GRPO (Group Relative Policy Optimization), the current standard for rubric-based RL with LLMs. Using Qwen3.5-4B as the policy model, RRT's macro criterion score across four benchmarks — Medical, Science, Rubrics as Rewards Science, and RubricBench — is 1.7 points above GRPO. The gains concentrate where they matter most: on hard and very hard criteria in Medical and Science, RRT gains 2.8 to 5.6 points over GRPO. This is not a uniform lift — it's a targeted improvement on the criteria that actually differentiate quality. The judge-budget result is arguably more practically important than the accuracy gain. Adaptive Fisher information selection — choosing which criteria to actually evaluate based on expected information gain, using a frozen RPN — lets RRT operate at half the criterion budget while staying within 0.1 points of full-budget GRPO. For anyone running rubric-based RL at scale, halving your LLM-judge API calls while matching baseline accuracy is real money. Integrity is mixed. The benchmarks (Medical, Science, RubricBench) are recognizable community datasets, not custom constructions. GRPO is the right baseline — it's genuinely current. But the evaluation is entirely self-reported: same team, same compute, no independent replication, no pre-registration. The monotonicity assumption — that all rubric criteria are monotone indicators of a single latent quality — is load-bearing and acknowledged but not stress-tested against adversarial or multi-dimensional rubrics. The obvious next experiment is scaling to larger policy models and multi-dimensional rubrics where the single-latent-trait assumption breaks. The authors likely stopped at Qwen3.5-4B because compute budgets for 70B+ RL runs are expensive, and because multi-dimensional IRT (MIRT) introduces identifiability headaches they'd need another paper to address. This is most likely option (a) — ran out of compute — with a side of (c), saving MIRT for the sequel.