Imagine you're playing a massive game of 'Name That Tune,' but instead of melodies you're matching sentence fragments to their origins across 31,170 Bible verses — and the author has scrambled the lyrics, translated them, and buried them inside Gothic fiction written a century ago. That's the retrieval problem this paper formalizes, and the structural insight is that the difficulty isn't uniform: direct quotations are easy (keyword overlap), allusions are brutally hard (meaning preserved, surface destroyed), and paraphrases sit in between. The question is whether dense neural encoders can close the gap where keyword methods go blind. The committed claim: a fine-tuned Danish sentence encoder (DFM-large) can retrieve biblical intertextual references in literary fiction at R@10 = 0.508, nearly doubling the zero-shot dense baseline and substantially outperforming BM25 on the hardest category — allusions — while a linguistically normalized BM25 remains a surprisingly strong zero-shot competitor. This is not a moonshot result; it's a careful proof-of-concept that computational retrieval can be a genuine scholarly tool for intertextuality, not just a toy. The benchmark itself is the load-bearing contribution. The authors draw on the critical commentary to a scholarly edition of Blixen's Seven Gothic Tales, constructing 189 annotated reference pairs and evaluating against all verses of historically plausible Danish Bible translations (both Old and New Testament). They then stratify references automatically by lexical overlap into quotations, paraphrases, and allusions — a clean move that lets you see exactly where each method succeeds and fails. BM25 with linguistic normalization (lemmatization, orthographic standardization) nails every quotation at R@10 but collapses on allusions (R@10 = 0.262). Dense models start weaker overall but hold up better on the semantic end. Fine-tuning uses hard negatives and five-fold cross-validation on a 189-instance dataset — tiny by ML standards, which is both the honest constraint and the methodological risk. The result: DFM-large's allusion R@10 jumps from 0.138 to 0.339, and overall R@10 climbs from 0.265 to 0.508. That's meaningful, but the absolute numbers tell you this is still a recall tool, not an oracle. Half the references don't appear in the top 10 even after fine-tuning. The most intellectually honest move in the paper is the false-positive audit. A literary scholar reviewed 30 rank-one predictions that the benchmark counted as wrong and judged 7 of them to be genuinely meaningful references the editorial commentary had missed. That's a 23% rehabilitation rate on 'errors,' which reframes the task: the model isn't just failing to find known references, it's surfacing plausible new ones. The authors resist the temptation to claim discovery and instead propose the framing of 'heuristic co-readers' — retrieval systems that generate candidates for expert-led close reading. The integrity setup is solid for a digital humanities paper: the benchmark comes from a published critical edition (not the authors' own annotations), the search space is exhaustive (all verses, not a cherry-picked subset), and the stratification is automatic rather than post-hoc. The main weakness is scale: 189 references from one author's one collection, with cross-validation on a dataset small enough that fold composition matters. Generalization to other authors, languages, or source corpora is entirely untested. The bigger picture: this paper participates in the growing argument about whether dense retrieval can replace or complement keyword search for tasks where surface form is deliberately disrupted — translation, paraphrase, allusion, historical language change. The answer here is 'complement, with fine-tuning': BM25 wins on quotations, dense models win on allusions, and the best practical system would likely combine both. For digital humanities, the real unlock isn't automation but triage — reducing a scholar's search space from 31,000 verses to 10, with a hit rate that's useful but not exhaustive.