Imagine you trained a sous chef by showing them how to "julienne carrots" a thousand times. They're perfect at it. Now you say "cut carrots into thin strips" and they stare at you blankly, or worse, start dicing onions. That's the core finding here: vision-language-action models (VLAs) — the hottest architecture for getting robots to follow natural language commands — are catastrophically sensitive to instruction phrasing, even when they're built on top of vision-language models (VLMs) that handle synonyms just fine. The language robustness doesn't survive the action-finetuning step. The numbers are genuinely alarming. π₀.₅ on the LIBERO benchmark hits 100% success for "switch on the stove" and 2% for "switch on the hot plate." That's not noise — that's a 98-point cliff from a synonym swap. Even a π₀ checkpoint finetuned with explicit rephrase augmentation still shows swings of up to 61 points. The authors run statistically tested single-edit experiments and an oracle phrase search showing that phrasing choice alone nearly closes a 21-point gap between in-distribution and out-of-distribution tasks. The problem isn't that VLAs can't do the task — it's that they've memorized specific word patterns rather than learning semantic task representations. The fix is elegant in its simplicity. Because the sensitivity is systematic (certain word choices consistently help or hurt), the authors treat it as a translation problem. They score many phrasings of a handful of training tasks, feed the scored evidence to a large language model, and have it distill 10-20 explicit rephrasing rules. At deployment, each incoming instruction gets rewritten once under these rules before the robot sees it. No retraining. No per-step verification. No weight changes. Results are solid. On the frozen π₀ policy, the rewrite layer improves success by 16-27% relative across twelve held-out tasks, with gains concentrated exactly where you'd want them — on out-of-distribution phrasings (adversarial, VLM-generated, human-generated). On LIBERO, in-finetune success climbs from 93.6% to 97.8%. The pipeline generalizes zero-shot to unseen tasks and instructions, and replicates across both π₀ and π₀.₅. The architectural insight is the important part. VLAs inherit visual grounding from their VLM backbones but lose linguistic robustness during action finetuning — likely because action datasets are tiny compared to VLM pretraining corpora, so the model overfits to the narrow phrasing distribution of the action labels. The rewrite layer is essentially a normalizer that maps the open distribution of human language back into the narrow dialect the robot actually learned. It's a shim, not a cure. The integrity picture is mixed. The benchmarks (LIBERO, Bridge V2 via π₀/π₀.₅) are community-standard, and the authors test across adversarial, VLM-generated, and human-generated phrasings — three distinct out-of-distribution regimes. Statistical testing of single-edit swings is a welcome rigor that most VLA papers skip. However, this is entirely the authors' own simulation pipeline, and the 10-20 rules are derived from the same task family they evaluate on, raising mild circularity concerns despite held-out task splits. The bigger question this paper forces is whether action finetuning fundamentally degrades language representations, or whether this is a data-scale problem that will vanish with larger action corpora. If the former, every VLA deployment needs a normalization layer — possibly forever. If the latter, this paper is a useful stopgap that will age out. The authors don't take a strong position, but the oracle phrase search results suggest the representation is still in there, just hard to reach, which leans toward the data-scale hypothesis.