Imagine you're moving into a new apartment with a box of assorted tools. You know you need to hang a shelf, but first you have to rummage through the box, pick the right drill bit, then walk to the wall and actually drive the screw. A toddler can pick the drill out of the box pretty reliably — but coordinating grip, aim, pressure, and walking over to the wall with the drill in hand? That's where things fall apart. This paper is about measuring exactly where humanoid robots fall apart in that same sequence. The committed claim: no existing benchmark jointly evaluates tool selection, tool manipulation, and mobile execution on a humanoid platform. HumanoidToolBench fills that gap with 18 tasks, 3 scenarios (tabletop, standing, mobile), 3 execution levels (selection only, manipulation, full mobile execution), and 2 tool-set modes (seen tools vs. unseen tools). The companion dataset, ToolBook, provides 3,100 demonstrations collected in MuJoCo simulation and on a physical Unitree G1 robot. The ladder here is more about benchmark design than beating a prior score. Existing tool-use benchmarks — RLBench, LIBERO, ManiSkill2 — focus on tabletop manipulation with fixed-base arms. They don't test locomotion, they don't test tool selection from a set, and they don't run on humanoid morphology. HumanoidToolBench is the first to combine all three in a single evaluation harness. The paper evaluates seven simulation policies (including ACT, Diffusion Policy, and NVIDIA's GR00T N1.7) and three policies on the real G1 robot, providing the first cross-policy comparison on humanoid tool use. Architecturally, the benchmark is agnostic — it evaluates policies from different families (behavior cloning, diffusion-based, transformer-based foundation models) rather than proposing a new one. The interesting architectural probe is on GR00T N1.7, NVIDIA's humanoid foundation model. The paper specifically stress-tests GR00T's language-conditioned tool selection: when given unseen tools, selection accuracy drops meaningfully, and when given unrelated language instructions ("pick up the hammer" when the task is about a screwdriver), the robot continues executing the original task rather than switching. This reveals a brittleness in how current foundation models bind language to action. Integrity is mixed but honest. The simulation environment is MuJoCo with the Unitree G1 URDF — standard, reproducible. Real-robot experiments are run but limited (3 policies, not 7), which the authors acknowledge as a hardware bottleneck. The ToolBook demonstrations were collected by the same team that designed the benchmark, so the training data and evaluation are not fully independent. Crucially, code and data are promised as open, which is the right move for a benchmark paper. But there's no external team validation yet, and the benchmark tasks were designed post-hoc to expose gaps, not pre-registered. The milestone to watch is straightforward: current policies show a steep performance cliff between selection-only and full mobile execution tasks. Closing that gap — getting a humanoid to select, grasp, walk, and use a tool at even 70% success across unseen tool sets — would signal that general-purpose humanoid tool use is approaching real utility. We're not there. The paper documents the gap rather than closing it, which is exactly what a benchmark paper should do. The obvious experiment not run: scaling ToolBook demonstrations by 10–50× and retraining the strongest policy (likely GR00T N1.7) to see whether the selection-to-execution gap is fundamentally a data problem or an architecture problem. The honest read is (a) — collecting 30k+ humanoid demonstrations is extremely expensive in both sim and real, and the team is at Seoul National University and Google, not a hardware-rich industrial lab. This is the experiment that NVIDIA or a well-funded robotics lab will run next.