Imagine you suspect a postal worker is slipping coded messages into otherwise normal-looking letters. You can't catch them by checking any single letter — the coded bits look like ordinary handwriting variation. But if you photocopy every letter before it goes out and compare the copies to what arrives, over enough letters, the pattern of tiny deviations adds up. That's this paper's core mechanism: replay prompts on a trusted server, compare token sequences, and accumulate evidence using e-processes — a statistical tool designed to never produce a false conviction no matter how long you keep watching. The committed claim: a prompt-level e-process can detect covert weight exfiltration through an LLM inference API with anytime-valid false-alarm control — meaning you can peek at the evidence after every response and stop the moment it's conclusive, without inflating your error rate. This matters because the alternative (hard per-token alarms) either fires too often on benign numerical noise or misses patient attackers who hide payload bits only in near-tie token decisions. The method sits in the family of sequential hypothesis testing, specifically e-processes (a generalization of likelihood ratios with game-theoretic roots). The architecture is lightweight: you maintain a trusted reference server, replay prompts, score whole-response mismatches against a calibration distribution learned from benign traffic, and multiply e-values across responses. The compute cost is essentially one extra inference per prompt — the statistical machinery itself is trivial. The key assumption is calibration transfer: that the mismatch distribution estimated on calibration traffic holds for future traffic. The paper evaluates on four models against two attack strategies. The seed-blind attack encodes bits regardless of token-probability ties, making it detectable but operationally simpler. The seed-aware attack is the interesting adversary — it only steals bits when the top two token probabilities are close enough that a flip looks like benign nondeterminism. This creates a fundamental channel-capacity-vs-detectability tradeoff: the stealthier the attack (smaller bandwidth), the harder to detect, but the slower the exfiltration. The e-process is evaluated against a hard per-token alarm baseline and shown to combine weak evidence across responses where the per-token alarm cannot. Integrity is reasonable for a workshop paper. The evaluation uses four named models and two attack types, which gives some breadth. The calibration-transfer assumption is stated explicitly — the authors know this is the load-bearing beam and say so. However, this is simulation-only: the trusted and compromised servers are both controlled by the authors. No independent red-team, no real-world deployment data. The benchmarks are custom (there's no community standard for weight-exfiltration detection), so all evaluation design choices were made by the same team making the claims. The milestone to watch: detection of seed-aware attacks at sub-1-bit-per-response bandwidth within 100 responses on production-scale models (70B+). The paper demonstrates the mechanism works; the open question is whether the calibration-transfer assumption holds in the wild, where prompt distributions shift, hardware varies, and attackers adapt. The gap between controlled simulation and real deployment is the chasm this line of work needs to cross. The obvious next experiment not run: a proper red-team exercise where a separate team designs the attack strategy and the detection team runs blind. The authors likely didn't run this because (a) it requires organizational infrastructure a workshop paper doesn't command, and (b) it's the natural next step they're saving for a full conference submission. The other missing piece is scaling — testing on 70B+ models where numerical nondeterminism is noisier and the attacker's hiding space is larger.