Imagine you run a counterfeiting lab, but instead of hiring expert engravers to design the security threads in banknotes, you hand the problem to three competing forgers and tell them: invent a mark that's invisible to the naked eye, survives a trip through the washing machine, and can be verified by any bank teller with a UV light. Then you grade their work automatically against every known counterfeiting technique. That's AutoMark: frontier LLMs acting as autonomous researchers to discover new watermarking schemes for LLM-generated text, evaluated by a rigorous automated pipeline that checks whether the marks actually hold up. The committed claim is straightforward: this is the first framework enabling fully autonomous discovery of distortion-free LLM watermarks that meet strict statistical reliability criteria. The authors don't just throw agents at the problem and eyeball the results. They build a three-layer evaluation harness — strict false-positive-rate verification, automated statistical testing of scheme correctness, and a composite ranking across detectability, quality, and robustness — that filters out schemes that look good on paper but would fail in deployment. Three frontier models (GPT-6 Astra, Opus 5, Gemini-3.8 Flash) collectively produce over 50 valid schemes, with several dominating all prior human-designed methods across every measured dimension. Where does this sit on the ladder? The baselines are the right ones: the current state-of-the-art in distortion-free watermarking, including schemes from Kirchenbauer et al. and their successors. The paper reports that discovered schemes outperform these baselines along all three axes simultaneously — not just trading robustness for quality or detectability for naturalness. That's the hard part of watermarking: improvements along one dimension typically degrade another. The agents found schemes that break this tradeoff, which is what makes the result non-trivial. Architecturally, this is an LLM-as-researcher setup: frontier models receive a specification of the watermarking problem (token-level logit manipulation under a distortion-free constraint), propose candidate schemes as code, and those schemes are executed and evaluated by the automated pipeline. The key structural insight is the evaluation harness, not the generation. Without the strict criteria — particularly the false positive rate verification and the multi-dimensional ranking — agent-generated schemes would be untrustworthy. The compute property the method leans on is inference-time reasoning from frontier models, not training. The agents are prompted, not fine-tuned. Integrity is mixed but mostly honest. The evaluation pipeline is automated and well-specified, with statistical tests for false positive rates and multi-dimensional ranking that resists single-metric gaming. Code is released. But the benchmarks are self-designed — there's no pre-registered protocol, and the ranking suite was built by the same team that built the framework. The manual study of discovered schemes (distilling key ideas into components, studying each component's impact) is a genuine effort at interpretability, but it's still the authors grading their agents' homework. No independent replication exists yet. The most interesting finding isn't the raw performance numbers — it's that the agents discovered fundamentally new ideas, not just recombinations of known techniques. The paper highlights one example: aligning watermark scores with a random per-request direction in logit space. This isn't a tweak to an existing method; it's a conceptually novel mechanism that a human researcher might have taken years to stumble on, or might never have tried. That's the real signal here: autonomous research producing genuinely novel algorithmic ideas, not just parameter sweeps. The obvious next experiment the authors didn't run is adversarial robustness testing against dedicated watermark-removal attacks beyond standard paraphrasing. The paper tests robustness against paraphrasing and text manipulation, but the arms race in watermarking is with sophisticated removal attacks (translation round-trips, targeted token substitution, generative paraphrasing with different models). The honest read: they likely ran out of scope in an already 52-page paper, and adversarial robustness is the natural sequel — probably being saved for the next paper or a dedicated red-teaming effort.