Imagine you're learning to drive with an instructor in the passenger seat. A bad instructor narrates every road sign. A good instructor stays silent through routine stretches and speaks only when you're about to miss something — merging without checking your mirror, drifting into the wrong lane. The hard part isn't knowing the rules of the road; it's knowing when silence is the right call and when a single sentence prevents a mistake. That's the core mechanism of EgoVoice: training a model to watch a continuous first-person video stream and decide, moment by moment, whether to stay quiet or deliver a spoken intervention. The committed claim is this: no existing system solves the joint problem of intervention timing and content generation from continuous egocentric multimodal streams, and EgoVoice is the first framework that trains and evaluates both together. The paper starts from HoloAssist, a dataset of real human instructors guiding people through physical tasks while wearing AR headsets. The authors perform source separation and speech resynthesis to extract clean instructor audio, then reformulate each session so the model must decide at every timestep: silence or speak. This is not a question-answering setup. The model never gets asked anything. Architecturally, EgoVoice fine-tunes an omni-modal large language model — one that ingests video frames, audio, and text jointly — using supervised fine-tuning on the cleaned HoloAssist data, followed by direct preference optimization (DPO) to sharpen the model's sense of when intervention is actually warranted versus when silence is better. The DPO stage is load-bearing: without it, models either talk too much (low precision) or almost never intervene (low recall). The paper reports experiments across both closed-source (GPT-4o) and open-source backbones, finding that zero-shot models — even powerful ones — rarely produce well-timed, meaningful proactive interventions. EgoVoice's fine-tuned system shows clear improvements in intervention timing, content relevance, and human preference over the zero-shot backbone. The ladder here is honest but thin. The baselines are zero-shot omni-modal LLMs prompted to behave proactively — there is no established SOTA for proactive egocentric spoken assistance because the task formulation itself is new. The paper is defining the benchmark rather than climbing an existing one. That's both a strength (genuine novelty in problem framing) and a weakness (we can't say 'beats X by Y%' against a mature competitor). Human evaluation shows EgoVoice preferred over zero-shot GPT-4o on timing and relevance, which is the right comparison given the field's maturity, but the absolute numbers suggest this is early-stage capability rather than deployment-ready performance. Integrity is mixed. The evaluation includes human preference studies (11 tables, 12 figures across 25 pages suggests thoroughness), and the data pipeline from HoloAssist is well-documented. But validation is same-team throughout — no independent replication, no pre-registration. The benchmark is self-constructed from the same dataset used for training, which creates circularity risk even though the train/test splits are presumably separate. The choice to evaluate against zero-shot baselines rather than against any fine-tuned competitor (even a naive one) leaves the ladder question partially unanswered. The real milestone isn't a single number — it's the transition from reactive to proactive AR assistance. Today's systems (Siri, Alexa, even multimodal assistants) wait to be asked. EgoVoice demonstrates that an omni-modal LLM can learn to initiate, but the gap between 'improved over zero-shot' and 'useful enough that a human wearing AR glasses actually benefits in real time' remains wide. The next concrete threshold: real-time inference on-device (or with tolerable cloud latency) during a physical task, evaluated by task completion metrics rather than preference ratings. That's probably 2-4 years out given current edge-compute constraints. The obvious experiment not run: deploying EgoVoice in a live, real-time AR loop where users actually perform tasks while wearing the glasses and receiving spoken guidance. The paper works entirely from recorded HoloAssist sessions — offline, post-hoc. My read: this is (a) infrastructure and IRB constraints. Building a real-time AR inference pipeline with human subjects is expensive and slow. The offline-first approach lets you iterate on the model before committing to that cost. But until someone closes that loop, we don't know whether well-timed interventions in hindsight translate to well-timed interventions in the moment.