Imagine you're a hiring manager who scores every résumé individually — a number from 0 to 1 for each candidate. That works fine when you're hiring for one role, but the scores don't transfer when you open a new position. What you actually want is a scoring rubric — a reusable function that can evaluate any résumé for any role. That's the shift TESS makes: from per-sample data weights to a learned selection network that generalizes across datasets and model scales. The committed claim: existing meta-learning objectives for training-data selection (specifically Meta-learning for Training-data Selection, or MTS) break when you replace per-sample weights with a neural selection network. The failure mode is specific and diagnosable — weight suppression and persistent reliance on easy-to-learn features — and the proposed fix is a new objective called Pointwise Value Matching (PVM) that avoids both pathologies. This matters because the field's current approach to data selection for LLM training is dominated by heuristics — perplexity scoring, deduplication, quality classifiers trained on proxy labels. MTS-family methods offer a principled alternative by learning data weights from a validation objective, but they've been stuck at the per-sample level, which means they don't transfer to unseen data. A selection network that generalizes would be genuinely useful at scale. The problem is that naively plugging a network into the existing MTS loss doesn't work. The diagnostic contribution here is arguably more valuable than the solution. The authors identify two specific failure modes: weight suppression (the network learns to assign near-zero weight to most samples, collapsing the effective training set) and easy-feature reliance (the network latches onto surface statistics rather than learning which samples actually help the target objective). These are not obvious failure modes — they emerge from the interaction between the meta-learning objective's gradient structure and the network's capacity to overfit. PVM addresses this by matching pointwise value estimates rather than optimizing the original MTS objective end-to-end through the selection network. The experiments cover LLM safety alignment and targeted instruction tuning, demonstrating transfer from subsets to full corpora and from smaller to larger models. The domains are well-chosen — safety alignment is a setting where data selection quality has outsized downstream impact. The transfer results are the load-bearing evidence. If PVM only worked on the same distribution it was trained on, it would be a marginal improvement. The claim that it transfers across dataset scale (subset → full corpus) and model scale (smaller → larger) is what makes this potentially useful for practitioners. However, the experiments are conducted at relatively modest scale compared to frontier LLM training, and the baselines compared are within the MTS family rather than against the full zoo of production data-selection heuristics. This is a well-scoped methods paper that identifies a real problem, diagnoses it clearly, and proposes a clean fix. It's not opening a new field — it's removing a specific bottleneck in an existing research program. The diagnostic insight about why the obvious approach fails is the part most likely to survive and influence subsequent work.