Imagine you're a restaurant health inspector, but every restaurant you visit has bribed the Yelp reviews to say "ignore the rats, give five stars." Your job isn't to eat the food — it's to spot which reviews are poisoned before they reach the diner. That's the prompt-injection detection problem in Retrieval-Augmented Generation: intercepting malicious instructions hidden inside retrieved documents before the LLM blindly follows them. RAG-PIBench is a new benchmark containing 4,876 contextual examples — a modest but carefully constructed dataset split into frozen train, validation, and a protected test set that researchers cannot peek at during development. The core contribution is methodological: the authors build a "leakage-aware" construction pipeline, meaning they actively prevent the kind of data contamination that inflates scores in benchmarks where train and test examples share phrasing, structure, or near-duplicate content. This sounds pedestrian until you remember how many ML benchmarks have been quietly invalidated by exactly this problem. The committed claim: leakage-aware benchmark design matters for reliable prompt-injection detection evaluation, and strong sparse baselines should not be ignored. This is not a claim about a new detection algorithm — it's a claim about how the field should measure progress. The paper tests keyword-based detectors, semantic-reference methods, TF-IDF with SVM and logistic regression, and transformer-based detectors (DistilBERT). DistilBERT takes the crown on the protected test set with F1 = 0.896 and PR-AUC = 0.968, but the surprise is that TF-IDF SVM and logistic regression remain competitive — a humbling result for anyone assuming transformers are the only game in town. The ladder here is instructive. The baselines are not straw men: keyword matching and semantic-reference methods represent what a security team might actually deploy today, while TF-IDF classifiers represent classical NLP that costs almost nothing to run. DistilBERT wins, but the margin over TF-IDF SVM is not embarrassing for the cheaper method. The paper is honest about this. What's missing is comparison against recent purpose-built prompt-injection detectors from industry (Lakera, Rebuff, Prompt Guard) — those would complete the ladder. Integrity is the paper's strongest suit. The frozen splits, protected test set, and leakage-aware pipeline are all deliberate design choices to prevent the score inflation that plagues detection benchmarks. Three figures and eight tables across 19 pages suggest the authors are showing their work rather than cherry-picking. However, the dataset is synthetic/constructed rather than drawn from real-world RAG deployments, which limits ecological validity. No code release or pre-registration is mentioned in the abstract. The milestone question is practical: at 4,876 examples and one architectural family (text classification), this benchmark needs to scale to tens of thousands of examples with diverse injection styles and multilingual coverage before it becomes the community standard for RAG security. The obvious next experiment — testing against actual production RAG pipelines with real adversarial injections rather than constructed examples — was not run. The honest read: ecological validity testing requires partnerships with RAG operators and access to real attack data, which an academic team likely lacks. That's a resource gap, not a methodological dodge.