Imagine you own a frozen yogurt machine with 200 flavors, but every customer who pulls the lever gets vanilla. The machine CAN produce mango-habanero or lavender-honey — the capability is in there — but the default path is vanilla because that's what the training signal rewarded most. DivLM is an attempt to recalibrate the lever so pulling it actually samples the full flavor space. The committed claim: a two-phase post-training pipeline — continued pre-training on creative writing plus RL with a custom diversity reward — can increase genre, tone, style, and named-entity diversity in LLM short stories by more than 9% on average versus alternative approaches, without degrading instruction following or output quality. This is not a new architecture. It's a training-regime intervention targeting the well-documented mode collapse problem in generative text. The method sits in the RLHF/RLAIF family, but replaces the usual human-preference reward with a composite signal measuring diversity across four narrative dimensions: genre, tone, style, and named entities. Phase one does continued pre-training on a creative writing corpus, then uses weight residuals (the delta between the fine-tuned and base weights) to restore instruction-following capability that continued pre-training tends to erode. Phase two applies reinforcement learning using the composite diversity reward. The two-phase design is the structural bet — separating domain knowledge injection from behavioral steering. The ladder question matters here. The paper reports >9% average improvement in diversity metrics against 'alternative approaches,' but the abstract does not name specific baselines with numbers. We don't know if the comparison is against vanilla sampling, nucleus sampling with different temperatures, or other diversity-promoting methods like DPO variants. The 9% figure is hard to contextualize without knowing what the floor and ceiling look like. Two LLM families were tested, which adds some generality, but without named models and named baselines, the ladder is underspecified. Integrity is mixed. The authors test on two model families, which is better than one, and they measure both diversity AND quality preservation — a dual-metric approach that reduces the risk of gaming one at the expense of the other. However, creative writing quality is notoriously hard to evaluate automatically, and the abstract doesn't specify whether human evaluation was involved. The diversity metrics themselves (genre, tone, style, named entity variation) are reasonable proxies but could reward surface-level variation without deeper narrative diversity. The milestone question for this line of work is whether diversity interventions can scale to novel-length generation and whether diversity gains transfer to downstream human preference judgments — not just automatic metrics. Right now we're at short stories with automatic diversity scoring. The next meaningful threshold would be human evaluators consistently preferring DivLM outputs over baseline in blind tests across extended narratives. The obvious experiment not run: human evaluation at scale. Automatic diversity metrics tell you the outputs are statistically different from each other, but not whether readers perceive them as meaningfully different or whether the diversity is shallow (swapping character names, shifting from 'dark' to 'somber') versus structural (genuinely different narrative arcs, conflict types, worldbuilding). The honest read: human eval is expensive and slow, and this is likely a compute/budget constraint rather than a negative result being hidden. The weight-residual trick for preserving instruction following after continued pre-training is the most reusable idea here — that's a technique other teams will borrow regardless of the diversity application.