You know how packing for a strict airline carry-on forces you to rethink every item — and sometimes the constraint makes you a better packer? You discover the shirt that does double duty, the toiletry bag that nests inside the shoe. You don't just shrink your suitcase; you reorganize what goes in it. That's the core mechanism here: quantization frees memory headroom, and the authors immediately spend that headroom on a small auxiliary module (LATTE) that actually improves navigation quality. The constraint was generative. The committed claim: a 4-bit quantized VLN model running on a 16 GB edge board can match or exceed its BF16 parent's navigation success rate, provided you co-design the runtime, the memory management, and a lightweight stop-action head together. This is not a compression paper that stops at 'accuracy held up okay.' It is a systems paper that profiles six quantization levels, seven stop-head variants, and multiple inference backends to find a deployable operating point — and then actually deploys it on hardware. The headline numbers tell the story crisply. The deployed IQ4NL model hits 58.02% success rate on R2R VLN-CE val-unseen, beating the BF16 baseline. It fits in 11.35 GB resident memory on a Jetson Orin NX 16 GB. Step latency drops 20.8× and step energy drops 13.3× versus storage-streamed BF16. But the most revealing number is the 36.8× energy variation across execution paths at the same 4-bit precision — meaning runtime selection matters as much as quantization itself. INT2, meanwhile, collapses entirely, establishing a hard floor. Architecturally, EdgeVLN sits in the quantized-LLM-on-edge family, specifically leveraging llama.cpp's GGUF quantization formats for a causal language model backbone (StreamVLN). The key structural choice is LATTE: a small causal transformer that reuses hidden states already computed by the backbone — no second vision encoder, no extra forward pass. It predicts a Stop Action verifier rank, improving the notoriously brittle stopping decision in VLN without adding meaningful compute. The whole stack runs through a custom llama.cpp VLN driver that reconstructs streaming context and prunes memory tokens on-board. Integrity is solid within scope. The evaluation covers all 1,839 R2R VLN-CE val-unseen episodes — a community benchmark, not a cherry-picked subset. Six backbone precisions and seven stop heads are ablated systematically, and the energy/latency/memory measurements come from the actual target hardware, not simulation. The limitation is that this is all one team, one board, one benchmark. There's no independent replication, no pre-registration, and the R2R benchmark, while standard, is a simulated indoor environment — real-world deployment with noisy sensors remains unaddressed. The milestone math is straightforward. At 11.35 GB on a 16 GB board, there's precious little headroom. The next meaningful target is running a comparable VLN model on a sub-8 GB edge device (Jetson Orin Nano or equivalent) — that would open deployment on drones and smaller mobile robots where even 16 GB is a luxury. Achieving this likely requires either 3-bit quantization that doesn't collapse (INT2 already fails here) or architectural distillation that reduces the backbone itself. The gap is probably 1-2 years if mixed-precision quantization research and smaller VLN backbones both advance. The obvious experiment not run: real-world deployment on an actual robot navigating a physical building. The R2R VLN-CE benchmark uses Habitat simulation with rendered panoramas. Sensor noise, lighting variation, actuator lag, and odometry drift are all absent. The honest read is (a) — they ran out of budget and hardware integration time. This is a systems-level contribution from a university lab, not an industry robotics team with a fleet of robots. The sim-to-real transfer paper is the obvious sequel.