Imagine you're a talent scout watching basketball tryouts through a foggy gym window. The fog blurs everyone equally — tall players still look taller than short ones, fast ones still look faster. Now someone hands you "fog-correcting" glasses. But the glasses don't just remove fog; they also strip away real height differences that happened to correlate with the fog pattern. You end up ranking players worse than you did squinting through the haze. That's the core mechanism of this paper: when proxy-metric and north-star-metric estimates share sampling noise because they're measured on the same customers, that shared noise preserves the rank ordering of candidate proxies better than you'd expect. The committed claim: in the regime where most experimentation platforms actually operate — hundreds of experiments, not thousands — subtracting the correlated sampling error between proxy and north-star effect estimates makes proxy-metric selection worse, not better, because the shared noise is rank-informative and the dominant sources of ranking error (noisy north star, small experiment archive) survive the correction untouched. This directly challenges recent work from major platforms (Microsoft's ExP team is the implied interlocutor) that treats shared sampling error as pure contamination to be removed. The method is clean and the validation is unusually well-layered. Provenzano uses a split-half design: estimate proxy and north-star effects on disjoint random halves of each experiment's customers, so shared sampling error literally cannot contaminate the comparison. This becomes the ground-truth benchmark. Then the paper measures how well the uncorrected (shared-error-included) ranking matches this split-half ranking. The Spearman correlation is 0.65 — the shared error carries genuine signal about proxy quality. When corrections remove more of the shared error, the ranking gets worse. Held-out experiments evaluated on disjoint halves reproduce the predicted ordering at ρ = 0.93. The archive is 262 experiments with 69 candidate proxy metrics — a real platform-scale dataset, not a toy. The simulation calibration is archive-matched: parameters are drawn from the empirical joint distribution, and the true ranking is known by construction. Even when simulations hand each correction method the true sampling covariance (an unrealistically generous gift), removal still degrades ranking in the small-archive regime. This rules out the escape hatch that correction merely fails because covariance is poorly estimated. Architecture-wise, this is classical econometrics and rank statistics, not ML. The tools are Spearman rank correlation, split-half estimation on disjoint customer samples, and calibrated Monte Carlo simulation. The key structural insight is that proxy selection is a ranking problem, not an estimation problem — a better point estimate of covariance doesn't guarantee a better rank ordering of candidates. This distinction is simple once stated but genuinely underappreciated in the experimentation literature. The crossover analysis is the most practically useful contribution. Correction does eventually pay off, but only when the experiment archive is large enough that north-star noise stops dominating the ranking variance. The paper maps this crossover boundary as a function of north-star noise and archive size, and provides three inexpensive diagnostic checks platform teams can run on their own data. This is the kind of actionable output that turns a theoretical insight into a deployable tool. The main limitation is scope: one archive, one platform. The 262-experiment, 69-proxy dataset is respectable but not enormous. The generalization claim rests on the simulation calibration being representative. The obvious next experiment — replicating on a second platform's archive — would dramatically strengthen the result.