Imagine you're assembling IKEA furniture while wearing oven mitts and looking through a keyhole. You can see fragments of the screw and bracket, but your own hand keeps blocking the view. What saves you is the pressure you feel through the mitt — you know whether you're gripping the screw head or the shaft, even when you can't see it. OccluDex is that pressure sense for robot hands, algorithmically fused with the fragmentary geometry the camera can still see. The committed claim: a hierarchical 3D visuo-tactile representation, pretrained on human demonstrations via multi-scale masked autoencoding, that beats all tested baselines on dexterous manipulation tasks where the robot's own hand occludes the object — by 12.6% on unseen geometries and 8.3% on seen ones. This isn't a new manipulation policy; it's a new perceptual backbone that makes downstream RL policies more robust because they start from a richer, occlusion-resistant state estimate. The architecture is a two-stage hierarchy. First, a multi-scale masked autoencoder ingests partial 3D point clouds at multiple resolutions, learning to reconstruct occluded geometry from fragments — essentially training the system to hallucinate what's behind the hand. Second, tactile contact tokens (from simulated BioTac-style sensors on the Shadow Hand's fingertips) are fused into the geometric features through cross-modal attention. The encoder is pretrained on synchronized human visuo-tactile demonstrations, then frozen as a perceptual backbone for downstream RL. This freeze-and-transfer strategy is the load-bearing architectural choice: it decouples perception pretraining from policy learning, which is why zero-shot sim-to-real transfer works at all. The evaluation covers two tasks: faucet rotation (one full clockwise revolution) and tabletop object reorientation (180-degree flip without toppling). The baselines include state-of-the-art methods, though the paper doesn't exhaustively name every competitor's architecture. The 12.6% and 8.3% margins are on success rate, which is the right metric for manipulation but hides variance. Physical experiments on a real Shadow Hand demonstrate zero-shot sim-to-real generalization on unseen objects — the strongest possible validation short of a multi-lab replication. Integrity is a mixed bag. The sim-to-real transfer on a physical Shadow Hand is genuinely impressive and non-trivial to fake, but the simulation benchmarks are author-designed tasks, not community-standard benchmarks like the DexGraspNet or DexArt suites. The baselines are described as 'strongest state-of-the-art' but are not all individually named with architecture-level detail in the abstract. The paper is 8 pages, submitted to IEEE, with no mention of code release or pre-registration. The milestone question is about scaling to multi-object, multi-step manipulation. Two tasks with known geometries (faucets, tabletop objects) are proof-of-concept. The next real milestone is robust in-hand manipulation of arbitrary household objects — think: a robot unloading a dishwasher where every plate, cup, and spatula is a novel geometry. That likely requires scaling from the current small set of unseen test objects (order of dozens) to hundreds of diverse geometries with varied surface properties, probably 2-3 years out if the sim-to-real gap continues closing at this rate. The obvious experiment not run: multi-finger, multi-object sequential manipulation in cluttered scenes — the dishwasher problem. The honest read is (a): the Shadow Hand hardware setup and simulation pipeline are expensive to extend to cluttered scenes, and two clean benchmark tasks are enough for an 8-page IEEE submission. The second missing experiment is ablating the tactile modality more aggressively — how much does performance degrade with fewer or noisier tactile sensors? That's likely saved for the journal version.