Imagine you're a copy editor working on a long magazine article, not just fixing individual sentences but restructuring whole sections — moving paragraphs, resolving pronouns that now point to nothing, ensuring the simplified version still reads as a coherent piece. That's document-level text simplification, and it's a fundamentally harder problem than sentence-level simplification because the unit of analysis is the entire discourse, not a single clause. This paper asks: can off-the-shelf multilingual LLMs do this for Estonian, a language where most NLP simplification research simply doesn't exist? The committed claim: Gemini-2.0 and LLaMA-3.3 can produce document-level Estonian simplifications with near-native fluency and strong meaning preservation, using carefully designed prompting strategies. Five models are tested — Gemini-2.0, LLaMA-3.3, GPT-4o, DeepSeek-R1, and Qwen-2.5 — across three prompting strategies: single-pass generation, a pipeline of modular agents (splitting the task into sub-steps), and a guideline-augmented pipeline. This isn't a new architecture paper; it's an evaluation and methodology paper that establishes the first baselines for a task-language combination that had none. The ladder here is unusual because there IS no prior art for document-level Estonian simplification. The authors are essentially constructing the first rung. They compare across their own five models and three prompting strategies, finding that Gemini-2.0 and LLaMA-3.3 significantly outperform GPT-4o, DeepSeek-R1, and Qwen-2.5. The weaker models show notable grammatical errors — a telling result given Estonian's complex morphology (14 cases, rich agglutination). But without an external established baseline system for this specific task, the paper is grading its own homework within its own cohort. The evaluation framework is the paper's most durable contribution. They combine automatic metrics for readability, semantic preservation, and discourse coherence — including novel document-level coherence metrics — with a structured manual annotation protocol. This dual-track evaluation is important because automatic metrics for morphologically rich languages are notoriously unreliable when used alone. The manual annotation protocol adds real signal, though the paper's 12-page length and 2 tables suggest the annotation scale is modest. Architecturally, this is prompt engineering over frozen multilingual LLMs — no fine-tuning, no adapter layers, no language-specific model surgery. The three prompting strategies represent increasing levels of task decomposition: single-pass (do everything at once), pipeline agents (break it into steps), and guideline-augmented pipelines (break it into steps with explicit simplification rules). The finding that pipeline approaches help is consistent with the broader LLM literature on chain-of-thought and task decomposition, but the specific interaction with Estonian morphology and discourse structure is new. The integrity picture is mixed. Resources are publicly available, which is excellent for a low-resource language paper. But the absence of a pre-existing benchmark means all evaluation design decisions were made by the same team reporting results. The manual annotation protocol helps, but we don't know annotator count, inter-annotator agreement details, or whether the evaluation framework was locked before results were observed. For a LREC workshop paper, this is par for the course — but it means the numbers need independent replication before they're load-bearing. The real value here is methodological: a template for how to bootstrap document-level simplification evaluation in any low-resource language. Estonian is the case study, but the prompting strategies, evaluation framework, and coherence metrics are portable. The next milestone isn't a bigger number — it's whether another team in another morphologically rich language (Finnish, Hungarian, Turkish) picks up this framework and confirms the patterns hold. If the pipeline-agent approach consistently outperforms single-pass for agglutinative languages, that's a genuinely useful finding for the ~200 low-resource languages where LLM-based simplification could matter.