Imagine you hire an intern who memorized every page of every tool's man page but has never actually typed a command into a terminal. You say "scan that subnet for open ports" and they confidently produce a command that looks plausible — right tool name, reasonable-looking flags — but they've swapped -sV for -sS, put the port range before the target, or invented a flag that doesn't exist. The command fails silently or, worse, does the wrong thing. That is the state of LLMs doing cybersecurity tool invocation today, and KaliBench is the exam that finally measures how bad it is. The committed claim: no existing benchmark directly measures whether LLMs can produce syntactically correct, executable CLI commands for real cybersecurity tools — and when you build one, every open-weight model fails more than it succeeds. KaliBench fills this gap with 8,504 natural-language-to-command pairs covering 1,642 tools across 23 capability dimensions and all 5 standard security phases (reconnaissance through post-exploitation). The dataset is constructed from tool manuals via a deterministic canonicalization pipeline, meaning evaluation is reproducible and alias-aware — if nmap accepts both --open and -open, both count. The ladder here is revealing. In the hardest setting — unrestricted, where the model must pick the right tool AND produce the right arguments with no hints — no open-weight model exceeds 42% exact-command accuracy. The paper tests 24 configurations across general-purpose and security-focused models, so this isn't one cherry-picked failure. The baseline isn't a straw man either: they compare across three evaluation modes of increasing difficulty (tool name given, tool name plus partial args given, nothing given), which lets you see exactly where models break. They break overwhelmingly at argument construction — choosing the right flags, binding them to the right values, getting the order right. The architectural contribution is the verification pipeline itself. A three-stage system — LLM-based semantic validation, sandboxed terminal execution, and human review — produces deterministic reward signals without needing a live target network at training time. This is the "runtime-free verifiable rewards" claim, and it matters because it means you can do reinforcement learning on CLI correctness without spinning up vulnerable VMs for every training step. They demonstrate this by fine-tuning an 8B parameter model using both supervised fine-tuning and RL with these verifiable rewards, pushing it to performance comparable with a 685B mixture-of-experts model. Integrity is solid for a benchmark paper. The dataset construction is transparent: manuscript-grounded extraction, deterministic canonicalization, alias-aware matching. The sandbox execution stage catches commands that parse correctly but wouldn't actually run. Human-in-the-loop refinement addresses edge cases. Code and data are released on GitHub. The main integrity gap is the usual one for benchmarks: the paper evaluates its own benchmark, and contamination risk in future model training against this public dataset is unaddressed. The practical milestone is clear. At 42% accuracy in the unrestricted setting, we're roughly halfway to a model you'd trust to draft commands without human review. The 8B-to-685B equivalence result suggests that targeted training data — not raw scale — is the bottleneck. If the community can push an 8B model from 42% to ~80% exact-match accuracy on this benchmark, you have a genuinely useful copilot for penetration testing workflows. That's probably 1-2 years of iteration on training recipes, not a hardware problem. The experiment the authors didn't run: end-to-end agentic evaluation where the model chains commands across a multi-step penetration test. KaliBench deliberately measures single-command accuracy in isolation. The obvious next step is sequential tool use — does getting one command right improve the next? — and likely they're either saving this for a follow-up or the evaluation infrastructure (sandboxed multi-step execution with ground-truth attack paths) is genuinely hard to build. My read: saving it for the next paper, because the single-command benchmark is already a substantial contribution and the multi-step setup is a different engineering challenge entirely.