Imagine you're translating a novel from Japanese to English, then asking someone to grade the English version. You'd expect the grade to depend on the novel's quality — but it also depends, heavily, on which translator you picked. Swap translators and the grade changes, sometimes drastically. Applied ML researchers do this every day with pretrained image encoders: they pick one (often whatever's popular or convenient), freeze it, extract features, and train a downstream model on those features. The encoder choice is treated like plumbing — invisible infrastructure. This paper argues it's more like choosing a translator for a courtroom deposition: the choice is consequential and should be audited. The committed claim: different pretrained encoders map the same images into different feature spaces, and those differences propagate into sharply different out-of-sample predictions. The authors call this representation risk and demonstrate it across six diverse tasks — hedonic house pricing, racehorse performance forecasting, breast-cancer histopathology, chest-radiograph pneumonia detection, continuous facial age estimation, and rice disease classification. Ten encoders spanning legacy (ResNet-50) to modern (SigLIP 2, DINOv2, CLIP, EVA-02) are compared under controlled conditions: common dimensionality reduction, identical heads, and group-safe data splits. The spread is not subtle. For house prices, SigLIP 2 reaches test R² = 0.629 versus ResNet-50's 0.396. For racehorses, DINOv2 hits R² = 0.105 versus ResNet-50's 0.029. No encoder dominates all tasks. The architecture here is deliberately simple, and that's the point. All encoders are frozen ViT-family or CNN-family models used as pure feature extractors. Downstream heads are either ridge regression or shallow neural networks — nothing exotic. The methodological contribution is the workflow, not a new model: benchmark multiple encoders, select on locked validation data, combine features only when validation evidence justifies the cost, and report split-level uncertainty. The authors implement this in an open-source package called LOOKAGAIN-ML. The simplicity is load-bearing: it means the gaps are attributable to the encoder, not to head-architecture confounds. Integrity is a mixed picture. On the positive side, the paper uses six genuinely diverse applied domains rather than cherry-picking a single favorable task. Training, validation, and test splits are locked, with group-safe partitioning (e.g., no data leakage across property-level or patient-level groups). Repeated random partitions are used for some tasks, revealing instability — the horse feature-union results, for example, don't hold up across splits. On the negative side, these are the authors' own tasks and splits, not community benchmarks. There's no pre-registration, and the code release status is unclear beyond the named package. The validation is honest about where gains are small (horses, rice), which is a good sign. The milestone question is more about workflow adoption than a single number. The paper doesn't claim a new SOTA on any established benchmark — it claims that the encoder-selection step itself is an underappreciated source of variance. The concrete gap is between current practice (pick one encoder, never look back) and the proposed practice (benchmark several, select on validation, report uncertainty). The next number to watch is whether encoder-selection variance gets reported in applied ML papers the way hyperparameter sensitivity already is. If five major applied-ML venues adopt encoder-comparison tables within three years, the workflow has landed. The obvious experiment not run: fine-tuning. The paper restricts itself to frozen encoders with at most limited adaptation (the 'limited adaptation' results are mentioned but not deeply explored). Full fine-tuning would likely compress the representation gaps — the question is by how much, and at what compute cost. The honest read is (a) compute budget: fine-tuning ten encoders across six tasks is expensive. But also (c) strategic: a frozen-encoder workflow is simpler, more reproducible, and more directly maps to the 'encoder as infrastructure' framing that makes the paper distinctive. Fine-tuning would muddy the message. What matters here is less the specific numbers and more the naming of a category of risk that applied researchers routinely ignore. If you use pretrained encoders as feature extractors — and an enormous amount of applied ML does exactly this — this paper is telling you that your results have a hidden dependency you're not reporting. The workflow is modest, the implementation is straightforward, and the contribution is primarily conceptual: making the invisible visible.