Imagine you're coaching a chess club, but instead of teaching strategy directly, you hand each player a rulebook, a stack of recorded games, and say 'figure it out — you have 10 practice matches.' Some players study their own losses, some watch what the top players did, and a few start rewriting their opening playbooks mid-tournament. That's Adversarial Heuristic Learning (AHL) — and this paper formalizes it, builds a competition arena for it, and discovers that the best AI coach still loses to seasoned humans in half the events. The committed claim: LLM-based agents can read game rules, analyze replays, and iteratively revise executable game-playing programs without any model fine-tuning — and this process is now benchmarkable across 12 diverse adversarial games with 1,920 real human programs as opponents. The authors call this paradigm Adversarial Heuristic Learning, distinguishing it from both classical reinforcement learning (which updates weights) and pure prompt-based reasoning (which doesn't iterate on code). The key constraint is that model weights stay frozen; all learning happens through code revision. The benchmark, AAArena, is the load-bearing contribution. Twelve authentic adversarial games — not toy problems — with an evaluation protocol modeled on real-world programming competitions. Agents interpret rules from specifications, choose opponents strategically, analyze match replays, and revise their game-playing agents to climb a ladder populated by archived human submissions. The evaluation enforces fixed match and evaluation budgets, preventing brute-force iteration. This is the first structured benchmark that tests the full loop: read rules → play → analyze → revise code → play again. The headline result: Claude 3.5 Opus paired with Claude Code earns 6 gold medals across the 12 games, meaning it tops the human ladder in half the competitions. But no evaluated configuration cracks the remaining 6 ladders. Performance degrades consistently as rule specification complexity increases — longer, more intricate rule documents correlate with weaker agent performance. This is a meaningful finding: the bottleneck isn't code generation or strategy refinement per se, but game comprehension from natural language specifications. The mechanistic findings are where practitioners should focus. Opponent selection matters — agents that strategically pick who to play against improve faster than those matching randomly. Dense feedback (detailed match logs rather than just win/loss signals) supports better policy revision. And agents learn from both on-policy replays (their own matches) and off-policy replays (watching other players' matches), which mirrors how human competitors study recorded games. These are actionable design choices for anyone building agentic code-revision loops. The integrity picture is mixed. The benchmark uses 1,920 real archived human programs — not synthetic opponents — which is strong. But the evaluation is entirely self-contained: the authors designed the benchmark, chose the games, and evaluated their own configurations. No independent replication exists. The model configurations tested are not exhaustive, and the specific games chosen could favor certain architectures. The 'completedmodels' placeholder in the abstract suggests the paper may still be in flux. The honest gap: they didn't test the agents against live human competitors adapting in real time, and they didn't evaluate whether the learned strategies generalize across games or are game-specific overfits. The first is likely a logistics issue — running live competitions is expensive and slow. The second is almost certainly being saved for follow-up work, because cross-game transfer would be the killer result that elevates AHL from benchmark curiosity to general capability claim.