Imagine a spell-checker that doesn't just flag misspelled words — it flags correctly spelled words that are wrong in context. "Duck" is fine in a hunting essay, dangerous in a cockpit transcript. That's the core mechanism here: the same physical motion (a punch, a throw, a kick) is safe or unsafe depending on what's in front of the robot. CSF is the contextual spell-checker for whole-body humanoid motion. The committed claim: a training-free safety filter that uses scene context — not just prompt inspection or geometric constraints — to decide whether a generated motion is safe. The method works by generating paired safe and unsafe reference trajectories from the same motion generator, then defining an affine safety value between them and enforcing it via a control barrier function quadratic program (CBF-QP). No retraining, no labeled motion datasets, no bespoke safety classifiers. The architecture choice is elegant in its simplicity. CSF sits downstream of any text-conditioned motion generator as a plug-in filter. It takes natural-language safety rules ("don't punch near a person"), uses the generator itself to produce reference trajectories for safe and unsafe versions, and then constructs a real-time CBF-QP that tracks the safe reference while provably avoiding the unsafe region. The key structural insight is that the generator's own output space defines the safety boundary — no external safety model needed. This is tested across four architecturally distinct generators (MDM, MoMask, MotionGPT, OmniH2O), which is a meaningful breadth claim. The numbers are strong for a first demonstration: CSF activates the intended safety rules in 100% of explicitly unsafe and scene-triggered unsafe cases across all four generators. Danger-event rate drops by up to 90%. Benign motion preservation sits at 88-100%, meaning the filter is not over-conservative. The real-world demonstration on a Unitree G1 humanoid is the integrity highlight — this is not just simulation. The robot physically prevented from executing unsafe motions in human-interaction scenarios. The ladder position is harder to assess cleanly because the authors are defining a new problem category rather than competing on an existing benchmark. Existing safeguards — prompt-level filtering, geometric constraints, labeled motion data — are acknowledged but not formally benchmarked head-to-head with controlled metrics. The paper argues these approaches don't address scene-dependent safety at all, which is fair but means the "90% reduction" is measured against unfiltered baselines, not against the best competing filter. Classical geometric constraint approaches would be the natural head-to-head, and that comparison is conspicuously absent. The milestone question is where to watch closely. The current system handles discrete rule activations (person detected → suppress punch). The real unlock is continuous, compositional safety reasoning — multiple overlapping rules, ambiguous scenes, partial occlusions, moving targets. The paper doesn't claim to solve this, and the gap between "four curated scenarios on a Unitree G1" and "robust safety in unstructured environments" is large. But the CBF-QP formalism is the right mathematical framework for composability, so the foundation is sound. The obvious experiment not run: adversarial prompt engineering against CSF. If the safety rules are natural-language, can a cleverly worded prompt fool the context detector while still producing dangerous motion? The authors likely didn't run this because (a) adversarial robustness is a separate paper and (b) the results might be uncomfortable for a first-demonstration paper. This is the experiment that will determine whether CSF is a real safety mechanism or a polite suggestion.