Imagine you discover a coupon-clipping service that saves you 20% at the grocery store. Great — until the store notices half its customers are clipping, raises shelf prices, and buries new fees in the fine print. Your coupon still works, but it's fighting a moving target. And the shopper behind you who doesn't clip? They're paying the new shelf price. That's the core mechanism of this paper, except the coupons are LLM-powered personal assistants, the store is a simulated rental market with six adaptive landlords, and the fine print is a mandatory fee and a pre-selected add-on that only an authorized assistant can cancel online. The committed claim: evaluating AI assistants at the individual level overstates their value because it ignores seller adaptation. When you freeze sellers in place, executing assistants save adopters $13.70 per renter-day. But when sellers are allowed to adapt — raising headline rents while cutting explicit fees — they claw back roughly a third of those gains, leaving $8.70. The paper's second claim is subtler and more important: the calibrated behavioral forecast predicted unassisted consumers would get hit with +$3.60 in spillover costs, but the actual simulation showed only +$0.40 (CI: -$0.6 to +$1.3). The behavioral model got the direction right but wildly overestimated the magnitude. The mismatch is the finding. Architecturally, this is an agent-based model (ABM) where LLMs play all roles — consumers, assistants, and sellers — with an algorithmic pricing tool guiding seller behavior. The simulation runs two mandates (advisory vs. executing) across paired branches: frozen sellers vs. adaptive sellers. Thirty market runs provide statistical power. The design borrows from computational economics but replaces hand-coded agent rules with language-model reasoning, which is what makes the seller adaptation emergent rather than scripted. This matters: the fee-setting behavior that determines how gains split between adopters and sellers isn't hard-coded, it's decided by the seller-side LLM responding to market conditions. The integrity picture is mixed but honest. The authors build an analytical benchmark and a behaviorally calibrated rule-based market as ex-ante predictions, then show where the LLM simulation diverges from both. That's good scientific practice — they're not just reporting wins. The paired frozen/adaptive branches are a clean causal identification strategy. But the entire validation regime is simulation-on-simulation: LLMs evaluating LLM-generated behavior. There's no empirical rental market data grounding the calibration, no independent replication, and the six-seller market is small enough that idiosyncratic LLM behavior could dominate. The seller-model swap and fee-transfer experiments partially address this, showing that fee conduct is seller-determined regardless of which LLM sits behind the seller, but the circularity concern remains. The milestone question here isn't about hardware but about ecological validity. The paper demonstrates the principle — market feedback erodes individual-level AI assistant gains — in a 30-run, 6-seller, ~50% adoption simulation. The next concrete threshold is demonstrating this in a market with hundreds of sellers and heterogeneous adoption rates, ideally calibrated against real transaction data from a platform like Airbnb or a rental aggregator. Until that happens, this is a well-constructed thought experiment, not an empirical finding. The obvious experiment the authors didn't run: varying the adoption rate continuously from 5% to 95% to map the full adaptation curve. At what adoption threshold do sellers start adapting aggressively? Is there a tipping point where non-user costs spike? The paper fixes adoption at 50%, which is clean for demonstration but misses the dynamics that matter for policy. My read: they're saving the adoption-rate sweep for the next paper, since the current framing already has enough moving parts and the 50% split makes the cleanest narrative.