Imagine you're a bouncer at a nightclub with a dress code. Most bouncers define one very specific outfit that counts as 'acceptable' — blue shirt, black pants, brown shoes. Most people showing up in perfectly fine outfits get turned away, and you waste everyone's time. Now imagine instead you define a neighborhood: anything within two changes of tonight's reference outfit is fine. Suddenly most well-dressed people walk right in. That's HammingMark. The core claim: existing semantic watermarking methods impose a fixed semantic preference on every generated sentence, but LLMs already have strong, context-dependent preferences about what sentence comes next. When the watermark's preference and the model's preference collide, the system has to resample repeatedly — sometimes dozens of times — to find a sentence that satisfies both. This resampling tax is what the authors call 'semantic narrowing,' and it degrades output quality on constrained tasks while burning compute. HammingMark replaces the rigid target with a Hamming ball in hash space centered on the previous sentence's semantic hash, so any candidate within a configurable edit distance counts as watermark-valid. The architecture is straightforward: each candidate sentence is encoded via a pretrained sentence embedder and then hashed into a compact binary code using locality-sensitive hashing. The watermark check is a Hamming distance comparison — does this candidate's hash fall within radius r of the previous sentence's hash? Because the hash is many-to-one, semantically diverse sentences can map to the same or nearby codes, preserving the model's natural output distribution. The dynamic center shifts sentence-by-sentence, making the watermark pattern hard for an attacker to predict or strip without access to the hash function. On the ladder: experiments on C4 (open-domain generation) and BookSum (long-form summarization) pit HammingMark against SemStamp and k-SemStamp, the current semantic watermarking baselines. HammingMark matches or exceeds their detection rates (z-scores and AUC) while requiring only 2.2 candidate samples per accepted sentence versus 8.1 for k-SemStamp — a 72.8% reduction. On BookSum, where semantic constraints are tighter, HammingMark achieves the highest detection rates with the highest or tied-highest ROUGE-L scores. The robustness tests include paraphrase attacks and sentence-level perturbations, where HammingMark holds up well. Integrity-wise, the validation is solid but self-contained. Benchmarks are standard community datasets (C4, BookSum), metrics are well-established (ROUGE-L, z-score, AUC, perplexity), and baselines are current named methods, not strawmen. However, there's no independent replication, no pre-registration, and no code release is mentioned. The hash function choice and LSH parameters introduce degrees of freedom that could be tuned post-hoc, though the authors test across multiple Hamming radii. The milestone that matters here isn't a single number — it's operational viability. At 2.2 samples per sentence, watermarking becomes cheap enough to deploy at inference time without meaningful latency cost. The next threshold is adversarial resilience against adaptive attacks: an adversary who knows the scheme exists (Kerckhoffs' principle) and specifically targets the hash-neighborhood structure. The paper tests standard paraphrase attacks but not adaptive ones designed to exploit Hamming-ball geometry. The experiment the authors didn't run: testing against an adaptive adversary who knows the watermark uses Hamming neighborhoods and specifically crafts perturbations to push sentences just outside the detection ball while preserving semantics. This is the obvious next paper. My read: they're saving it — the defensive extension (tighter hash functions, randomized radii) is a natural follow-up, and demonstrating the base scheme first is the rational publication strategy.