Imagine you're learning to cook by watching a chef, but the chef only tells you whether the final dish tastes good — never whether you chopped the onions wrong or burned the garlic. You'd eventually learn, but it would be agonizingly slow and you'd repeat the same mistakes hundreds of times. That's the state of online reinforcement learning for computer-use agents (CUAs) today: the agent clicks through a desktop environment, and only at the very end does it learn whether the whole trajectory succeeded or failed. ComputerSD is a kitchen timer that goes off at every step. The committed claim: combining token-level on-policy self-distillation (OPSD) with trajectory-level GRPO, regulated by a real-time GUI analyzer that scores each individual action, produces a meaningfully better computer-use agent than outcome-only RL. On OSWorld-Verified, ComputerSD achieves gains of 1.9 percentage points on the general-purpose Qwen3-VL-8B-Thinking backbone and 4.1 percentage points on the specialized EvoCUA-8B backbone versus GRPO alone. The architecture is a two-signal hybrid. A fine-tuned GUI analyzer watches each screen transition after every agent action and produces two things: a natural-language guidance string (privileged context about what just happened on screen) and a step-level value score. The guidance feeds into OPSD — a method where a teacher model with access to privileged information rescores the student's token probabilities to provide dense learning signal. The value score acts as a gate: if the step was bad, the OPSD signal is suppressed so the agent doesn't learn to imitate its own mistakes with higher confidence. Meanwhile, the standard trajectory-level GRPO reward still operates on complete rollouts. Both signals are optimized jointly in a fully asynchronous training framework. The two core engineering insights here address real failure modes. First, static teacher guidance can drift out of alignment with what the student policy is actually doing — the student takes a weird action, and the teacher's pre-computed advice becomes irrelevant. By generating guidance from the actual executed GUI transition rather than a pre-planned trajectory, ComputerSD keeps guidance grounded in reality. Second, OPSD can create probability shifts that are locally correct at the token level but globally wrong at the step level — the teacher might upweight tokens for an action that looked linguistically reasonable but was functionally incorrect. The step-level value score from the GUI analyzer filters these out. The baseline comparison is honest but narrow. GRPO is the right comparison — it's the dominant online RL method for these agents — and the gains are consistent across two different backbones, which reduces the chance of architecture-specific overfitting. The out-of-distribution evaluation adds credibility. However, the absolute numbers remain modest: we're talking about agents that still fail on the majority of computer tasks, and the improvements, while statistically meaningful, are single-digit percentage point deltas. The paper doesn't compare against other dense-reward approaches outside the self-distillation family. The validation relies entirely on OSWorld-Verified, which is a community benchmark — good — but there's no pre-registration, no independent replication, and the GUI analyzer itself is trained by the same team, which introduces a circularity concern: how much of the gain comes from the analyzer's quality versus the training method? An ablation separating these would strengthen the claim considerably. What matters for the field is the general principle, not the specific numbers. The idea that you can extract dense, step-level supervision from environment feedback (screen transitions) and use it to regulate token-level distillation is transferable well beyond OSWorld. If this approach scales — particularly with stronger GUI analyzers or multimodal critics — it could significantly compress the sample complexity of online CUA training. The next milestone to watch: whether this method or its descendants can push OSWorld-Verified success rates past 50%, where computer-use agents start becoming practically useful rather than research demonstrations.