Imagine you walk the same route to work every day. Some days it's sunny, some days it's dark and raining, sometimes you're jogging, sometimes dawdling. You recognize the same landmarks each time — the cracked sidewalk, the red mailbox, the tree that leans left. Now imagine you had to explain to someone ELSE how two separate GoPro recordings of that walk relate to each other, matching landmarks across different lighting, speeds, and camera angles. That's the task EgoGears sets for multimodal large language models, and every single one chokes on it. The committed claim: current MLLMs can answer questions about what happens within a single egocentric video reasonably well, but they fail systematically when asked to carry that understanding across multiple independent recordings of the same physical environment. The mean accuracy drop is 22.5 percentage points — not a gentle degradation, a cliff. This isn't about video length or format boundaries; it's about a specific cognitive bottleneck the authors call observation-evidence binding and ordered route-state tracking. The benchmark itself is carefully constructed. 126 human-collected egocentric recordings cover 39 outdoor routes, with repeated traversals under varying movement speeds and lighting conditions. From these, the authors derive 567 single-video questions (testing local perception of visual, spatial, and motion evidence) and 1,487 multi-video questions (testing whether evidence stays bound to the correct observation and composes into consistent route relationships). 531 of the multi-video questions explicitly require alignment across independent recordings. The single-video split serves as a control: it measures the raw perceptual signal available to the model, so the multi-video drop can be attributed to transfer failures rather than perception failures. The evaluation is broad: 29 single-video and 31 multi-video MLLM configurations across six model families in the main leaderboard. Among the 20 configurations evaluated comparably on both splits, every model performs worse on multi-video questions. The authors are careful to control for confounds — they hold answer format and scoring fixed, and show the gap isn't simply explained by longer inputs or recording boundaries. The bottleneck is structural, not incidental. The strongest prior baselines here are the current crop of frontier MLLMs — GPT-4o, Gemini, Claude, and open models like LLaVA and InternVL families. EgoGears doesn't claim to beat any of them; it claims to expose a failure mode none of them have solved. The ladder question is diagnostic, not competitive: the benchmark reveals that models which look capable on single-video VQA are systematically fragile when the same knowledge must persist across encounter boundaries. No model family escapes the drop. The integrity picture is solid for a benchmark paper. The dataset is human-collected (not synthetic), the evaluation covers six model families with multiple configurations each, the code and benchmark are publicly released, and the experimental design controls for obvious confounds. The main limitation is that this is a first-party evaluation — no independent group has yet replicated the findings on their own hardware or with their own scoring pipeline. The question selection methodology (how the 2,054 questions were derived from 126 recordings) could introduce subtle biases, though the controlled single/multi-video split design mitigates the worst risks. The real value here is diagnostic clarity. The field has been measuring video understanding with aggregate accuracy on single-video benchmarks, which conflates perceptual failures with transfer failures. EgoGears separates these, and the separation reveals that the transfer problem is much harder than the perception problem. For anyone building embodied AI systems — robots, AR glasses, autonomous navigation — this is the gap that matters. A robot that understands what it sees right now but can't connect that understanding to what it saw yesterday on the same route is functionally amnesiac.