Imagine you're trying to thread a needle, but you can only look at it from across the room through a single security camera. You know how to thread needles — your hands are fine — but the camera angle hides whether the thread is actually entering the eye. A friend walks over, holds up their phone at a better angle, and streams you a close-up. Suddenly you nail it. Your skill didn't change. Your observability did. That's SpatialHarness. The committed claim: frontier multimodal models like GPT-6 Astra already possess fine-manipulation policy capability, but fail because fixed physical cameras don't reveal the task-critical spatial relationships (plug-to-socket alignment, block-on-peg centering). SpatialHarness provides complementary virtual views at test time — no retraining, no new physical sensors — and this alone dramatically improves success rates. The paper frames this as an observability problem, not a policy problem, which is a meaningful reframing for the robotics community. The mechanism works in three stages. First, the system builds and maintains an online simulated scene synchronized with the real robot's workspace using what the authors call interaction-aware scene synchronization — distinguishing objects as static, held, or in transition so the digital twin doesn't drift during contact-rich manipulation. Second, it identifies which spatial relationships are task-critical (e.g., the coaxial alignment of a plug and socket). Third, it renders virtual camera views from angles that expose those relationships and feeds them to the frozen foundation model policy alongside real camera images. The results are striking on the four real-robot tasks tested. Plug insertion jumps from 26.7% to 66.7%. Tower of Hanoi goes from 0% to 100%. These aren't simulation-only numbers — they're on physical hardware with a real robot arm. The frozen GPT-6 Astra policy is identical in both conditions; only the visual input changes. That's a clean experimental design that isolates the contribution. The integrity picture is mixed. On the positive side, these are real-robot experiments, not sim-only validation, and the baseline is the same frontier model without the harness — a fair comparison. On the negative side, four tasks is a small evaluation suite, the task selection could favor the method (all are geometrically precise tasks where viewpoint obviously matters), and there's no comparison against alternative observability solutions like additional physical cameras, active vision policies, or tactile sensing. The 0% → 100% result on Hanoi is eye-catching but should raise calibration alarms — perfect scores on small N trials can be artifacts of sample size. Architecturally, SpatialHarness sits in the test-time augmentation family — cousin to chain-of-thought prompting and retrieval-augmented generation, but in the embodied domain. It treats the foundation model as a frozen black box and wraps it with a spatial reasoning layer. The sim-sync component leans on standard pose estimation and object tracking, while the view selection is the novel piece. The compute overhead is a simulation render loop, which is lightweight compared to policy fine-tuning. The big question this paper doesn't answer is scalability to cluttered, deformable, or novel-object scenes where building an accurate digital twin is itself the hard problem. The four tasks use known, rigid objects with clean geometry. The honest read on why they didn't test messier scenarios: building reliable sim-sync for deformable objects or cluttered scenes is a separate research problem, and they're scoping this paper to prove the observability thesis cleanly. Fair enough — but the real unlock is whether this generalizes beyond geometric puzzles to the messy manipulation tasks that actually block deployment.