Imagine you're a driving instructor giving road tests. You don't need every student to drive the same 50 routes — after watching someone blow through three stop signs, you already know they're not ready for the highway merge. You pick the next maneuver based on what you've already seen. That's the core mechanism here: Item Response Theory (IRT), borrowed from educational psychometrics, applied to cybersecurity evaluation of coding agents. The committed claim: SecProbe is an adaptive evaluation framework that combines IRT-based ability estimation with on-demand synthesis of repository-scale vulnerability-repair tasks, achieving comparable ability estimates to exhaustive testing while requiring agents to solve up to 29.5% fewer tasks. This is not a new vulnerability dataset — it's a new way to administer the test. The benchmark itself is substantial: 353 tasks spanning six programming languages and 151 CWE types, tested against nine frontier models using two agent harnesses. The headline number is brutal — the best success rate peaks at 28.33%. That means even the strongest coding agent fails roughly three out of four vulnerability-repair tasks. This isn't cherry-picked difficulty; the CWE coverage is broad and the languages diverse. The gap between where these agents are and where they'd need to be for autonomous security remediation is enormous. Architecturally, SecProbe sits at the intersection of psychometrics and program synthesis. The IRT backbone treats each task as an 'item' with estimated difficulty and discrimination parameters, and each agent as a 'test-taker' with a latent ability score. When the framework needs more evidence in a particular difficulty range, it can synthesize new tasks on the fly rather than relying on a fixed pool. This is the key structural innovation — the benchmark isn't static, it grows where measurement is weakest. The agent harnesses (two unspecified configurations) wrap frontier LLMs to interact with repository-scale codebases, not just isolated code snippets. The integrity picture is mixed. The 353 tasks and 151 CWE types provide breadth, and comparing against random and one-shot baselines is reasonable for establishing adaptive efficiency. But the baselines are selection-strategy baselines (random vs. adaptive ordering), not competing evaluation frameworks. There's no comparison against other adaptive testing systems or established security benchmarks like CyberSecEval or the NYU CTF benchmark. The nine frontier models are unnamed in the abstract, which makes independent verification harder. The 29.5% efficiency gain is meaningful but modest — it's a convenience improvement, not a paradigm shift. The milestone question is where this gets interesting. At 28.33% peak success, we're far from agents that can autonomously find and fix vulnerabilities. The useful tracking number is that success rate: when it crosses ~70% on a benchmark of this breadth, you're looking at agents that could plausibly assist human security teams at speed. We're roughly a 2.5× improvement away from that threshold, which at current scaling rates could be 2-3 years — but the history of LLM capability on structured code tasks suggests gains will be uneven across CWE categories. The obvious experiment not run: testing whether SecProbe's synthesized tasks actually match the difficulty distribution of real-world vulnerabilities found in production codebases. The authors have a synthesis pipeline and an IRT model, but the ecological validity question — do these generated tasks predict performance on actual CVEs discovered in the wild? — remains open. My read: this is being saved for the follow-up paper, where they'd need access to curated real-world vulnerability corpora and probably industry partnerships to validate.