Imagine you're training a new line cook by showing them cooking videos. Some videos are shot from the chef's perspective at their station — same counter height, same knife angles. Others are GoPro clips from food festivals: entertaining, broadly relevant, but the camera's too high, the tools are different, and half the time you can't see the hands. Both are "cooking data," but the first batch teaches knife technique in two days; the second takes two weeks and the cook still fumbles plating. Ego4WAM is a systematic study of exactly which properties of that first-person video matter when the student is a robot arm, not a line cook. The committed claim: egocentric human data's value for robot policy learning is NOT a single-axis scaling story (more hours = better robots). Instead, four properties — human-robot alignment, task diversity, data duration, and supervision type — interact in non-obvious ways that the field has been conflating. The paper disentangles these under a fixed world-action model backbone so the comparisons are clean. The most consequential finding is about alignment. When the camera viewpoint, hand morphology, and workspace geometry match between human demonstrator and robot, out-of-distribution generalization improves substantially and the amount of target-task robot data required drops. This isn't surprising in retrospect, but the field has been scaling ego-video quantity without controlling for this variable. Duration and task diversity have different downstream signatures: more hours help world modeling but task diversity helps policy breadth. These are not the same axis and shouldn't be traded off carelessly. A genuinely useful result for practitioners: video-only data (no action labels) still provides a strong foundation for subsequent video-action fine-tuning. This matters because action-labeled human data is expensive to collect and requires instrumented setups. If you can pre-train on cheap unlabeled ego-video and then fine-tune on a smaller labeled set, the data pipeline economics shift significantly. The architecture family is a world-action model — a jointly trained vision-language-action backbone that predicts both future visual states and motor commands. The paper fixes this backbone across all experiments, which is the right move: it lets you attribute performance differences to data properties rather than model differences. Validation comes through closed-loop policy evaluation on real robot hardware and the RoboDojo simulation benchmark. The integrity story is mixed. Real-robot evaluation is the gold standard and they do it, but the paper is fundamentally an ablation study, which means the design space of comparisons is author-chosen. There's no pre-registration, no community benchmark for "data property ablation," and the specific axes tested (alignment vs. duration vs. diversity vs. supervision) are reasonable but not exhaustive. Domain randomization, augmentation strategies, and embodiment transfer gaps are not ablated. The successor question is straightforward: cross-embodiment transfer at scale. The paper studies one robot morphology; the obvious next experiment is whether aligned ego-video from humans transfers differently to a bi-manual setup vs. a mobile manipulator vs. a humanoid. The authors almost certainly plan this — it's the natural next paper — and the fixed-backbone design makes it modular enough to extend. The honest read is (c): saving it for the next paper, because this one is already dense with ablations.