Imagine you've been studying for a geography exam using a general atlas. Then someone hands you a detailed regional map of one specific country. You'll obviously get better at naming rivers in that country — but the real question is whether you can now navigate there. This paper hands a Russian BERT model a regional map of legislative language and measures whether it got better at the naming part. It did. The navigation part remains untested. The committed claim: continued pretraining of a Russian ModernBERT encoder on ~194M tokens of Russian legislative text produces a model (RuModernBERT-ruLaw) that achieves lower masked-token cross-entropy on held-out court-decision segments. The cross-entropy drops are 0.109 nats at 512 tokens, 0.071 at 2,048 tokens, and 0.066 at 8,192 tokens. This is a standard domain-adaptation result — not a first, but a careful application to a language and domain where such work is thin. The baseline comparison is minimal but honest: the paper compares only the original RuModernBERT against the adapted version. There is no comparison against other Russian legal encoders, no multilingual legal BERT baselines, and no classical TF-IDF or n-gram language model to sanity-check the magnitude of the gain. The authors deserve credit for flagging exactly this limitation. The ~0.07-0.11 nat improvement is real but its practical significance is left as an open question — we don't know how much of that translates to better retrieval, summarization, or classification on actual legal tasks. The NER evaluation is where the integrity picture gets complicated. Both models score entity-level F1 above 0.998. Sounds impressive until you read the fine print: 99.95% of test spans share the same normalized surface form and class as spans in the training split. The authors explicitly state this provides 'limited evidence about transfer to previously unseen forms.' That's unusually honest self-reporting, but it also means the NER evaluation adds nearly zero information about the adapted model's capabilities. Architecturally, this is straightforward encoder-side domain adaptation using ModernBERT's sliding-window attention at extended context lengths (up to 8,192 tokens). The corpus contains 304,382 legislative documents. The paper explains its overlapping-window strategy, averaging rules for positions covered by multiple windows, and exact entity-boundary scoring with commendable clarity — the 11 figures and editable diagrams are a pedagogical strength. The paper's real contribution is methodological transparency. The authors distinguish corpus tokens from tokenizer positions, mark their confidence intervals as sensitivity-to-masking rather than population uncertainty, and refuse to overclaim on the NER result. This is a model of honest reporting in a subfield where inflated claims are routine. But methodological honesty does not substitute for a missing downstream evaluation — the paper stops exactly where it would start to become useful to a practitioner deciding whether to deploy this model. The trajectory here is clear: the obvious next step is a downstream legal task evaluation — document classification, clause extraction, legal question answering, or judgment prediction on Russian court data. The authors almost certainly know this. The most likely read is that this is a first-paper release establishing the adapted model, with task-specific evaluations planned for a follow-up. The model exists; the evidence for using it doesn't yet.