Imagine you're a bartender watching a busy room. You don't predict where someone will walk next by clocking their current speed and direction — you watch their eyes. Are they scanning the bar menu? Looking at a friend across the room? Reaching for their coat? The eyes telegraph intent seconds before the feet follow. This paper asks the same question for robots: what information channels actually help predict where a person will go next in a building? The committed claim: this is the first systematic ablation study comparing inertial (body motion), occupancy (scene geometry), semantic (room labels and object categories), and intent (eye gaze fixation + explicit task goals) signals for indoor human motion prediction using a diffusion model. The headline result is a 42% improvement over a constant velocity baseline, but the real finding is the ranking of information sources. Explicit intent dominates everything else. Eye gaze alone provides spatial information that partially substitutes for an occupancy map. Semantic labels help, but less than they do in outdoor navigation — indoors, the ambiguity of destinations within a single building overwhelms coarse room-type labels. The dataset is modest but carefully constructed: nine naive subjects wearing Meta Aria glasses, navigating ten real university buildings while performing simulated daily tasks — going to class, getting coffee, visiting offices. That's 238 minutes of trajectory data with synchronized IMU, eye tracking, scene understanding, and ground-truth position. The diffusion model architecture is a conditional denoising framework that takes various subsets of these input channels and predicts future 2D position trajectories. The ablation design is clean: each information source is toggled on and off to isolate its marginal contribution. The ladder here is honest but limited. The primary baseline is constant velocity prediction — the simplest possible extrapolation. The 42% improvement is meaningful but the paper doesn't compare against the strongest trajectory prediction models from the outdoor/autonomous driving literature (Social Force, Trajectron++, AgentFormer). The authors acknowledge this is a different domain — indoor spaces with sharp turns, doors, and elevator decisions — but the absence of a learned baseline beyond their own architecture means we're measuring relative contribution of input channels, not absolute SOTA performance. The integrity picture is mixed. Nine subjects is small. The buildings are all from one university campus, introducing institutional layout bias. The "simulated daily activities" are scripted task sequences, not truly naturalistic behavior — subjects were told to go to specific rooms. Eye gaze fixation data from Meta Aria glasses is a specific hardware dependency. The ablation design itself is the strongest methodological feature: systematic toggling of input channels with consistent architecture is the right way to answer "what matters" questions. Code is promised upon acceptance but not yet available. The finding that matters most for robotics practitioners is the eye gaze result. Gaze fixation was especially useful for predicting deceleration — the moment someone is about to stop or turn. This is exactly the hardest prediction regime for constant-velocity models and the most safety-critical for robot path planning. The paper also shows that gaze partially substitutes for occupancy maps, meaning a wearable sensor could provide spatial intent signals without requiring pre-mapped environments. The 20-year trajectory of this work points toward intent-aware human-robot cohabitation. If indoor robots — delivery bots, assistive systems, warehouse cobots — need to predict human paths, this paper argues they should invest in intent estimation (gaze tracking, task context) over ever-finer scene reconstruction. The gap between "scene understanding helps" and "intent understanding transforms" is the core actionable insight.