Imagine you're a food critic who rates restaurants by comparing exactly two dishes — say, the carbonara versus the bolognese. You declare the kitchen "Italian-biased" based on which scores higher. But if you'd picked a different carbonara recipe, the ranking flips. Your conclusion wasn't about the kitchen's bias; it was about which specific sentence you happened to write down. That's the state of LLM stereotype evaluation today, and this paper is the first to name the problem precisely and offer a structural fix. The committed claim: single minimal pair comparisons — where you pit "Women are nurturing" against "Men are nurturing" and check which gets higher log-likelihood — are logically unreliable because swapping the stereotyped attribute ("nurturing" → "caring") can reverse the model's apparent preference. This isn't a minor instability. It means a substantial fraction of published bias measurements are artifacts of word choice, not measurements of the model's associative structure. The fix is a dual minimal pair setup. Instead of one axis of comparison (group A vs group B for a fixed attribute), you add a second axis: the stereotyped attribute vs an alternative attribute, for a fixed group. This creates a 2×2 cell where consistency can be checked. If the model prefers "Women are nurturing" over "Men are nurturing" AND prefers "Women are nurturing" over "Women are aggressive," you have converging evidence. If only one comparison holds, the original measurement was noise. The authors then introduce a mutual information (MI) metric that quantifies the strength of association between social groups and attributes, which aggregates cleanly across languages and models — a genuine methodological contribution. The data-augmentation framework is applied across English, Russian, Spanish, and Chinese stereotypes, filling gaps in existing datasets like CrowS-Pairs and StereoSet by generating paraphrases and alternate attributes. This multilingual scope matters: bias benchmarks have been overwhelmingly English, and cross-linguistic comparison has been essentially impossible without a commensurable metric. The MI-based metric is designed precisely for this commensurability. On the integrity front, the paper is honest about what it's doing: this is a measurement-methodology contribution, not a debiasing intervention. The authors don't claim to fix models — they claim to fix how we evaluate them. The validation is necessarily somewhat circular (you're measuring the metric's behavior, not an external ground truth), but the logical argument for dual pairs is sound on first principles. Code is public on GitHub. The ladder question is interesting. The baselines here aren't competing methods but competing metrics: CrowS-Pairs' pairwise comparison, StereoSet's scoring, and log-likelihood ratios as used in dozens of papers. The paper argues these are all vulnerable to the single-pair instability it identifies. It doesn't claim to produce different bias rankings — it claims that existing rankings are unreliable and its rankings are more robust. This is a harder claim to falsify but a necessary one to make. The successor experiment that's missing is the obvious one: applying the dual minimal pair framework at scale to re-evaluate published bias findings and showing how many flip. The paper demonstrates the instability exists but doesn't systematically quantify the false-positive rate of existing benchmarks. My read: this is being saved for a follow-up paper, because that quantification would be the headline result and it's too valuable to bury in a methods contribution.