Imagine you hand twelve architecture firms a plot of land, a budget, and six hours. No blueprints, no zoning guidance — just soil samples. The surprising result isn't that some build functional houses. It's that two groups independently reinvent load-bearing designs that match competing structural engineering philosophies neither was told about. That's what SciExam for ENSO does with AI agents and climate science. The benchmark gives agents real observational data for El Niño–Southern Oscillation, the planet's dominant mode of year-to-year climate variability, and asks them to build low-order stochastic models from scratch. Agents process observations, write their own diagnostic code (which is then frozen), and iterate on model development using only those self-authored diagnostics as feedback. Hidden graders then score the models on three axes: reproducing ENSO's statistical properties, recovering unobserved variables, and forecasting held-out years. The same graders score a published reference model on identical criteria. Six of twelve agent systems outscore the published model, primarily through better reconstruction and forecasting. The genuinely striking finding isn't the leaderboard position — it's what the agents build. When the authors simplify the strongest models' structures, each one turns out to be compatible with one of the two competing explanations for ENSO's warm-cold asymmetry, a live debate in climate science that the benchmark task never mentions. The agents aren't just curve-fitting; they're producing structures that map onto distinct physical hypotheses. Whether this reflects genuine scientific reasoning or sophisticated pattern matching that converges on similar functional forms is the load-bearing question. Integrity controls matter here. The authors run controlled experiments on the top-performing system under varied information conditions to test whether high scores come from memorizing the observational record. The evidence suggests they don't — model structure changes in response to the information provided, and performance degrades in expected ways when key data is withheld. This isn't airtight proof against memorization, but it's a meaningful step beyond 'we ran it and it worked.' The architectural landscape is diverse: twelve agent systems, each with its own code-generation and self-evaluation pipeline, building stochastic dynamical models. The key constraint is the frozen-diagnostics design — agents must commit to their evaluation criteria before they can optimize against them, preventing the obvious failure mode of agents gaming their own metrics. This is a clever structural choice that makes the benchmark meaningfully harder than 'write code that passes tests you also wrote.' The published reference model is a real baseline, not a straw man. It's a specific low-order stochastic ENSO model from the literature, scored by the same hidden graders on the same metrics. Six of twelve agents beat it, which means six also didn't — a distribution that suggests the task is genuinely difficult rather than trivially solvable. The benchmark tests something most AI-for-science evaluations dodge: can agents produce valid scientific models when there's no known correct answer to grade against? What this paper opens is a methodology for evaluating AI scientific research that doesn't require a ground truth. If the approach generalizes beyond ENSO — to other climate modes, to other fields with competing theoretical frameworks — it could change how we assess whether AI agents are doing science or performing it. The immediate limitation is that we're watching agents rediscover structures humans already know about. The real test comes when an agent proposes a model structure that doesn't map to any existing theory and turns out to be right.