Imagine a restaurant kitchen with three chefs who each have their own style — one favors bold spices, one favors acid, one favors fat. You'd expect a varied menu. But now hand them the same recipe card written in slightly different fonts and watch: one chef reads the card and decides to cook, two others read it and sit out. The dish that reaches the table isn't determined by the chefs' skills — it's determined by which chef the formatting activated. That's the selection mechanism this paper dissects, except the chefs are LLM families and the recipe cards are news presentations. The committed claim: model diversity among LLM trading agents is an illusion of the population, not a property of the order book. Li and Bandyopadhyay build synthetic markets populated by a fixed mixture of three language-model families — Qwen, Mistral, and a third — and feed them identical economic events (a financing announcement, a workforce reduction) packaged in different presentation formats. The headline finding is that presentation alone shifts which model family dominates submitted orders by 40-48 percentage points. In the financing event, Qwen's share of submitted orders swings by 48 pp depending on how the news is framed. In the workforce-reduction event, Mistral's share swings by 40 pp and the net order direction actually reverses sign. Crucially, this happens with negligible change in aggregate order counts — the total volume barely moves, but the composition of who's trading flips dramatically. The architecture is deliberately simple: three LLM families receive identical information in varied presentation bundles, each independently decide whether to submit a buy or sell order, and the resulting order flow is analyzed for composition shifts. This is not a reinforcement-learning market simulator or a multi-agent bargaining framework. It's closer to a controlled A/B test on LLM decision-making, using the market framing as a legible outcome variable. The analytical contribution is a decomposition showing why selection into trading — the act of some models choosing to sit out — can either improve or worsen price accuracy even when aggregate demand sensitivity stays constant. The mechanism is compositional: if the models that opt in happen to share a directional bias, opposing flow vanishes and the price signal degrades. The integrity picture is mixed. On one hand, the authors are unusually forthright about what they haven't done: they specify prospectively that cleaner replication and a known-value validation (where the true asset value is known, so price accuracy can be measured against ground truth) are needed. On the other hand, the current validation is entirely synthetic — LLMs trading in a constructed environment against each other, with no real market data, no pre-registration, and no independent replication. The presentation bundles themselves are a potential confound: we don't know if the effects are stable across prompt paraphrases, temperature settings, or model versions. The 44-page length suggests thorough exposition, but the 3-figure count hints at a primarily analytical rather than empirical paper. Where this lands on the ladder: there's a growing literature on LLM agents in market simulations (papers from groups at MIT, Stanford, and elsewhere using GPT-family models as traders), but most of that work focuses on whether LLMs can trade profitably or mimic human behavioral biases. The specific question of how presentation framing induces selection effects that alter order-flow composition is genuinely underexplored. The paper doesn't claim to beat any prior baseline on a performance metric — it claims to identify a mechanism that prior work hasn't isolated. That's a different kind of contribution, more conceptual than empirical. The 20-year question this paper points toward is serious: as LLM-based agents proliferate in real financial markets, will apparent model diversity translate into genuine diversity of opinion, or will correlated responses to information framing create hidden monocultures in order flow? If three model families collapse to one active family depending on how Bloomberg or Reuters formats a headline, the systemic risk implications are non-trivial. The paper is early-stage and synthetic, but it's asking the right question at the right time. The obvious next experiment the authors did not run — and they essentially admit this — is the known-value validation, where you assign the asset a true fundamental value and measure whether the selection-induced compositional shifts actually degrade or improve pricing relative to that ground truth. My read: this is a compute-and-scope limitation (option a), not a buried negative result. The decomposition is analytical, showing the mechanism can go either way; the empirical confirmation of which direction dominates in realistic settings would require substantially more infrastructure.