Imagine you're dropped into a foreign city with a phrasebook and a bicycle. The phrasebook gets you through common tourist routes — turn left at the cathedral, right at the fountain. But when someone asks you to 'go find the building with the weird mural and check if the café inside is open,' the phrasebook is useless. You need to improvise: ride somewhere plausible, look around, adjust, try again, and — critically — write down what worked so tomorrow you don't repeat the same dead ends. ASENA is a robot that does exactly this: it has a phrasebook (a learned 4B-parameter vision-language navigation policy called ASENA-VLN) AND a general-purpose coding agent that can write new programs on the fly, debug them against recorded failures, and save working solutions as reusable skills. The committed claim is twofold. First, ASENA-VLN alone sets new state-of-the-art on the two hardest vision-language navigation benchmarks: 68.7% success on Room-to-Room (R2R) and 70.2% on Room-across-Room (RxR), surpassing prior methods that used larger models or privileged inputs. Second, when you wrap that policy inside a coding agent that can inspect its own failures, retry with modified programs, and accumulate notes across episodes, performance climbs dramatically: from 72% to 98% on R2R and 65% to 89% on RxR over just ten self-evolution passes — with model weights completely frozen throughout. The agent isn't learning in the gradient-descent sense; it's learning in the software-engineering sense, shipping patches to its own codebase. Architecturally, ASENA-VLN sits in the vision-language-action (VLA) family: a shared decoder ingests RGB images and language instructions, then outputs body-frame waypoint trajectories. The clever bit is training on three heterogeneous tasks simultaneously — route-following instructions, visual question answering, and a new dataset of 'atomic navigation tasks' derived from geometric primitives (turn 30°, advance 2m, orient toward object). This multi-task recipe gives the policy both macro-route capability and fine-grained maneuvering. On top of this sits the coding agent layer: a general-purpose LLM that treats navigation, perception, and actuation as callable APIs, writes Python programs to compose them, inspects execution logs when things break, and commits working skills to a persistent workspace. The system runs on standard GPU hardware with a 4B-parameter policy — not frontier-scale compute. The ladder comparison is solid but has texture. On standalone VLN benchmarks, ASENA-VLN beats NaviLLM, MapGPT, and other recent methods by meaningful margins on both R2R (68.7% vs prior ~65%) and RxR (70.2% vs prior ~62%). The agentic benchmarks are newer and less standardized — the self-evolution results (72→98%, 65→89%) are measured on recurring 100-task subsets, which is a favorable setup since the agent literally accumulates task-specific knowledge. The authors are transparent about this: they call it 'persistent workspace evolution,' not generalization. On embodied QA, they report state-of-the-art accuracy with fewer interaction steps, though the comparison set is thinner. Integrity is mixed-to-strong. The R2R and RxR benchmarks are community-standard with public test servers, which is good. The self-evolution numbers come from a self-selected 100-task recurring subset, which is informative but not the same as a held-out generalization test. Real-world demos on a Unitree G1 humanoid show qualitative transfer — search, visual inspection, spatial reasoning, synthesized gestures without a pre-built map — but these are demonstrations, not controlled experiments. Code and project page are provided. No pre-registration, no independent replication yet. The milestone question is where this gets interesting for the field. The 98% on recurring tasks shows the ceiling when self-correction is allowed. The real unlock is whether this transfers: can ten passes of self-evolution on Task Set A improve performance on unseen Task Set B? If the accumulated skills generalize across environments (not just across repeated runs of the same tasks), you have something closer to genuine robot autonomy. The gap between 68.7% zero-shot and 98% self-evolved is 29 points — the question is how much of that gap survives when you remove task recurrence. The obvious experiment not run: testing self-evolution on a fully held-out task distribution where no task repeats. The authors surely know this is the killer experiment. Two honest reads: (a) simulator budget — running 10 full passes over novel tasks is expensive and the current recurring-task setup already makes the point about the mechanism, or (c) they're saving the generalization story for the next paper, since the current framing ('persistent workspace evolution') neatly scopes the claim without requiring it. Either way, that experiment is coming, and the result will determine whether ASENA is a clever engineering system or a genuine paradigm for robot self-improvement.