Imagine you're a wine judge at a blind tasting. Someone whispers 'the left glass won gold last year' before you sip. Suddenly your scores shift — not on a few borderline calls, but on more than half of every pair you taste. That's what's happening when we ask LLMs to judge whether an AI-generated idea is novel. This paper measures how badly, and the answer is: catastrophically. The core claim is blunt: LLM-based novelty evaluation, the mechanism increasingly used to validate automated ideation systems, is unstable to the point of meaninglessness under small design perturbations. The authors build a controlled evaluation set from OpenReview by mining reviewer comments that explicitly affirm or dispute originality, keeping only submissions with unanimous reviewer agreement at the extremes. They pair these vetted papers with vanilla LLM-generated ideas and then test six different judge configurations. The methodology is elegant — it isolates the judge from the generator and controls for the quality signal. The damage report is stark. Simply telling the judge which idea reviewers found novel — a framing cue, not new evidence — shifts pairwise accuracy by over 50 points on identical idea pairs. That's not noise; that's the entire measurement instrument bending to suggestion. The same framing change that helps one judge actively hurts another, meaning there's no universal 'fix the prompt' solution. This is not a calibration problem — it's a structural fragility in how LLMs process evaluative instructions. Retrieval-augmented generation and larger reasoning budgets — the usual engineering reflexes when an LLM underperforms — help little. Two purpose-built novelty evaluators, systems specifically designed and marketed for this task, are outperformed by the cheapest prompted baseline the authors construct. This is a credibility problem for every paper that has reported novelty gains using these tools without disclosing which judge, which prompt, and which framing was used. The architecture story here is straightforward: all six judges are prompted LLMs (the paper sits squarely in the prompt-sensitivity / LLM-as-judge literature), and the evaluation set is constructed via automated mining of OpenReview metadata paired with LLM-generated ideas as foils. The compute demands are modest — the signal is methodological, not scaling-dependent. What matters is the experimental design: controlled pairs, unanimous reviewer agreement as ground truth, and systematic variation of a single design choice at a time. The integrity regime is stronger than most LLM evaluation papers. Ground truth comes from an external source (OpenReview reviewer consensus at the extremes), not from the authors' own labels. The evaluation set construction is automated and reproducible. The main weakness is that reviewer unanimity at the extremes, while a strong filter, still inherits whatever biases peer review carries — reviewers themselves may anchor on framing cues in submissions. But this is acknowledged, and the alternative (no ground truth at all) is worse. The implications land directly on the growing industry of AI-powered research ideation tools. If the novelty measurement is this fragile, then reported improvements in idea novelty from systems like those built on top of GPT-4 or Claude are not trustworthy without disclosure of the full evaluation pipeline. The paper doesn't just identify a bug — it identifies a structural audit gap in how the field validates its most exciting claims.