You know that feeling when you're stuck on a home renovation and your experienced neighbor walks in, glances at the problem, and says 'You need a Simpson Strong-Tie — aisle 14, third shelf'? They're not smarter than you. They just have a mental index built from years of seeing what works where. That's the skill this paper tries to measure in AI: given a half-formed research question, can a system point you to the buried prior work that would actually change your approach? ScholarCatalyst is the first retrieval benchmark built from the ground up around author-confirmed inspiration. The team recruited 184 lead authors of 207 recent CS papers and asked them to label which earlier works genuinely advanced their projects, each with a written rationale. This isn't citation mining — it's asking 'what did you actually need to read?' The resulting dataset captures a signal that citation graphs miss entirely: the difference between perfunctory references and the paper that made the project possible. The headline result is sobering. Standard dense-embedding retrieval hits 0.48 Recall@20 on these queries. Agentic search — the kind that iteratively reformulates queries and calls the same retriever as a tool — manages only 0.42. Even an agent built on Claude Fable 5.1, a model that likely encountered the finished papers during pretraining, reaches just 0.51. The gap between what current systems achieve and what a knowledgeable human scientist does routinely is large and not closing through scale or tool-use alone. The architecture story matters. The paper tests both single-shot embedding retrieval (dense vector similarity against a corpus filtered to papers available before the project started) and multi-step agentic pipelines that can reformulate queries, filter by date, and iteratively refine. The agentic approach uses the embedding retriever as a sub-tool, so it has strictly more capability — yet it performs worse. This suggests the bottleneck isn't retrieval mechanics but the upstream ability to formulate the right query from an ill-defined research question. Integrity is a genuine strength here. The temporal filtering is rigorous: candidate papers are restricted to those published before each project began, preventing leakage from the finished work. Author annotations include detailed rationales, not just binary labels. The benchmark includes both 'gold' papers the authors actually used and 'silver' papers they confirmed could have helped but didn't find in time. The main vulnerability is scope: 207 papers, all in computer science, all from authors willing to participate — a self-selected and field-limited sample. The milestone question is where things get interesting. Current best is 0.51 R@20. A system that reliably hits 0.80 R@20 on this benchmark would represent something qualitatively new: an AI research assistant that genuinely surfaces non-obvious prior work. That's not just a retrieval improvement — it's a capability that changes how science gets done. The gap from 0.51 to 0.80 likely requires training recipes that go beyond retrieval fine-tuning and into something resembling domain intuition. The obvious missing experiment is fine-tuning a retriever or agent specifically on the ScholarCatalyst training split. The paper evaluates off-the-shelf systems and general-purpose agents but doesn't train on its own data. The most charitable read: they're establishing baselines first and the fine-tuning paper is next. The less charitable read: fine-tuning might not help much, which would be an even more interesting result to report.