Imagine you're a short-order cook. You handle raw chicken, then grab a spatula, then use that spatula on a salad plate. You've just created a contamination chain — and the fix isn't just washing your hands, it's tracing back every object you've touched since the chicken and cleaning each one in the right order, at the right cost, given what the restaurant can afford. That's the core mechanism of HygieneRoboBench: not "can a robot clean?" but "can a planner reconstruct a contact history and figure out the cheapest safe path forward?" The committed claim: no existing robotics benchmark jointly tests a planner's ability to (a) reconstruct contamination chains from execution history involving dual grippers and shared objects, (b) plan safe continuations after unexpected new contacts, and (c) do so while minimizing cost under user-specified priorities. HygieneRoboBench fills that gap with 624 instances across 134 task families, and a new hybrid planner called Hygiene-NSP that combines LLM grounding with constraint programming to hit 94.4% safe resolution and 90.4% optimal safe resolution. The benchmark design is the real contribution. Tasks encode contamination propagation through two grippers and shared objects, treatment costs (washing, sanitizing, replacing), and user priority profiles that create genuine tradeoffs. The evaluation separates controlled history comparisons from independent plan evaluation, so you can see whether a planner merely avoids contamination or actually finds the cheapest safe path. This matters because the paper's core finding is that safely completing a task does NOT guarantee lowest execution cost — a planner can be safe but wasteful. Hygiene-NSP itself is a three-stage pipeline: an LLM grounds natural-language task descriptions into structured representations, a contact-history reconstructor traces contamination chains backward through the execution log, and Google's CP-SAT solver jointly optimizes treatment and task execution under the user's priority constraints. The architecture is honest about what LLMs are good at (grounding ambiguous language) and what they're bad at (precise constraint satisfaction), offloading the hard combinatorics to a proven solver. The ladder is informative but narrow. Baselines include several LLM-based planners (GPT-4o, Claude, Gemini variants used as direct planners) and a symbolic PDDL planner. LLM-only planners struggle badly — they can identify contamination risks in isolation but fail to jointly optimize safety and cost. The PDDL planner handles safety but can't ground natural-language priorities. Hygiene-NSP's 94.4% safe resolution rate meaningfully beats all baselines on the full 624-instance set. But note: the baselines are all zero-shot or lightly prompted — no fine-tuned models, no retrieval-augmented planning baselines. The integrity picture is mixed in the usual ways. The benchmark is new and authored by the same team that built the planner, so the planner had home-court advantage. The evaluation protocol — separating history/profile/event comparisons from plan evaluation — is well-designed and reduces circularity. But there's no independent validation: no other team has run their planners on this benchmark yet, and the task families, while diverse, are synthetic household scenarios, not real-robot demonstrations. The milestone question is where this gets practical. Today: 624 synthetic planning instances with two grippers. The obvious next threshold is real-robot deployment where sensor noise, partial observability, and physical contamination (not just symbolic contact flags) enter the picture. The gap between symbolic contamination tracking and actual microbial transfer modeling is enormous, and the paper doesn't pretend otherwise. The successor experiment everyone will want — running Hygiene-NSP on a physical dual-arm manipulator in a real kitchen — almost certainly wasn't run because it requires a hardware setup and bio-safety validation pipeline that's a different research program entirely.