Imagine you're a locksmith who both picks locks and installs them. Someone hands you a lock you've never seen before — no YouTube tutorials, no forums, no prior art. Can you crack it? Can you fix it so nobody else cracks it? Now imagine you do fix it, and your client asks: "But is it really fixed, or did you just block the one trick you know?" That's the paper. Heldt, Turk, Landolt, and Fritz do something the AI-security community has talked about but rarely executed honestly: they test frontier AI models on vulnerabilities that have never been publicly disclosed. Five nonpublic software environments, including bugs the authors themselves privately reported while they remained unpatched. This matters because every public benchmark — CTF challenges, disclosed CVEs — suffers from contamination. If GPT-4 or Claude has ingested a write-up of a vulnerability during training, its "exploit generation" score is measuring recall, not reasoning. By using nonpublic bugs, this paper forces the models to generalize. The headline result is nuanced and that's what makes it credible. Repair scores beat attack scores in two of the five nonpublic environments, but attack scores beat repair in the other three. There's no clean "AI helps defenders more" or "AI helps attackers more" narrative here. The variation across vulnerability types is substantial, which means aggregate benchmarks that average across bug classes are hiding the signal. A model might be excellent at patching buffer overflows but useless against logic bugs, or vice versa. The most striking number: 92 out of 524 test intervals where the initial exploit was stopped still fell to a subsequent different exploit. That's a 17.6% failure rate on the "is the fix actually robust?" question. Passing an initial security test is necessary but nowhere near sufficient. The authors frame this as motivation for "subsequent resistance tests" — essentially, you can't declare victory after blocking one attack path; you have to probe whether the repair introduced new attack surface or left adjacent vulnerabilities open. Methodologically, the paper makes a deliberate choice that deserves respect: deterministic graders, not LLM judges. In a field where "GPT-4 evaluates GPT-4's security output" is disturbingly common, using researcher-developed scoring rubrics with deterministic pass/fail criteria removes a layer of circularity. The comparison set of 209 disclosed vulnerabilities and cryptographic challenges provides a contamination-controlled contrast, though the nonpublic sample of five environments is small enough that environment-specific confounders could dominate. The paper is short — 8 pages, 1 figure — and reads more like a position paper with empirical backing than a full experimental study. The sample size (five nonpublic environments) limits the statistical power of any per-vulnerability-type conclusion. But the contribution is primarily methodological: here is how you should benchmark AI cyber capabilities, and here is what changes when you do it honestly. The field's current reliance on public CTF benchmarks for release decisions is the real target. For practitioners making AI deployment decisions, the actionable takeaway is the subsequent-attack finding. If your red-team evaluation stops after confirming the AI-generated patch blocks the known exploit, you're missing roughly one-in-five cases where a different exploit still works. Vulnerability-specific attack-repair comparisons with iterative resistance testing should become standard protocol.