Imagine you're at a buffet with 200 dishes and a friend who knows your taste. Your friend doesn't need to be a chef — they just need to pick the right plate from what's already on the table. That's the core insight here: personalized generation isn't a capability problem (the LLM can already produce good candidates), it's a selection problem. A small, well-calibrated ranker that knows YOUR preferences beats a massive generalist reward model that knows everybody's preferences vaguely. The committed claim: a million-parameter MLP ranking model, trained on fine-grained user preference data and parasitically reusing the base LLM's internal embeddings, outperforms billion-parameter reward models across all nine tested personalization datasets — at four orders of magnitude lower scoring latency. This is accepted at NeurIPS 2026, which means the reviewers bought it. The paper starts with an empirical observation that should change how you think about personalized alignment. There exists a massive performance gap between what a base LLM generates on average and the best candidate hiding inside a large pool of its own outputs. Standard alignment collapses all users into one monolithic preference function. Best-of-N sampling could exploit the gap, but current reward models are both too expensive (billions of parameters scoring hundreds of candidates per query) and poorly calibrated for individual preferences — they optimize for aggregate human judgment, not yours. The authors call this 'untapped headroom' and demonstrate it empirically before proposing their fix. The architecture is deliberately boring, which is its strength. Take the base generator's internal embeddings — the representations it already computes during generation — and feed them into a lightweight multi-layer perceptron trained to rank candidates according to per-user preference data. No new encoder, no auxiliary model, no distillation pipeline. The MLP piggybacks on compute the generator already spent. The key structural choice: factorized ranking rather than pointwise scoring. The model learns to order candidates relative to each other for a specific user, not to assign absolute quality scores. This is what makes it work with small parameter budgets — relative ordering is a simpler function to learn than absolute valuation. The ladder comparison is where this gets interesting. The paper benchmarks against billion-parameter generalist reward models (the standard RLHF/DPO-style alignment scorers) across three personalized generation settings on nine datasets. The million-parameter ranker wins on every single dataset. That's not a marginal improvement on a cherry-picked benchmark — it's a clean sweep with a model that's 250× smaller. The scoring latency gap (four orders of magnitude) means you could score a pool of 1,000 candidates in the time it takes a reward model to score one. The paper also shows the ranker can guide generation itself (not just post-hoc selection), reducing the cost of materializing the N candidates in Best-of-N. The integrity story is solid but not ironclad. Nine datasets across three settings is broad coverage, and the comparison to billion-parameter reward models is the right baseline to pick. The NeurIPS acceptance adds a layer of peer review. However, the abstract doesn't mention pre-registration or code release, and independent replication is naturally absent at this stage. The personalized preference data scaling is the load-bearing assumption — if fine-grained per-user preference data is hard to collect in practice, the approach's real-world applicability narrows. The milestone to track: can this framework maintain its advantage when the candidate pool grows past 1,000 and when applied to open-domain generation rather than structured personalization benchmarks? The 10,000× latency advantage creates room to scale N dramatically, but ranking quality at N=10,000 with diverse generation is untested. The obvious experiment the authors didn't run — and this reads like a 'saving it for the next paper' situation — is deploying this as a live personalization layer on a production LLM with real user feedback loops, where preference distributions shift over time and cold-start users have no history.