Imagine you're grading a student's chemistry lab report by checking whether they mentioned the right reagent names — but never checking whether they actually ran the experiment. That's what keyword-matching benchmarks do for tool-use in small language models: they credit the model for saying the right words in the right neighborhood, even if the model never produced a structurally valid tool call. This paper catches such a false positive in the wild and builds a cheap diagnostic ladder to prevent it. The committed claim: lenient metrics like BLEU-4 cannot distinguish between a model that genuinely performs tool use and one that merely generates text that resembles tool calls. The authors demonstrate this with a matched-architecture pair — a 661.6M parameter model (65% code/technical training, no dedicated supervised fine-tuning for tools) and a 1,109M model (web-heavy curriculum with 6B-token tool-SFT). Both share decoder, tokenizer, and special tokens. On BLEU-4, they score 0.660 vs 0.650 — essentially identical. But verbatim-reproduction checks on training examples separate them completely: the 600M emits valid tool calls with generalized arguments on 6/6 examples; the 1B does so on 0/6 across multiple checkpoints. The failure mechanism is precise and locatable. A first-token probe shows the 1B assigns probability 10⁻⁴ to 10⁻⁵ to the <|toolcall|> special token — the trigger was effectively erased by the web-heavy training phase. This is not a subtle architectural mismatch; it is a missing prior that prevents the model from ever entering the tool-call format. The 600M, with its code-heavy pretraining, retained that prior without any dedicated tool SFT. This is the core insight: pretraining distribution matters more than labeled tool-use data if the labeled data gets drowned out. The repair recipe is surprisingly cheap. A targeted SFT phase using a diverse corpus, 5× higher learning rate, and just 2,202 steps (~3.3 GPU-hours) raises the 1B's valid emission rate from 0.100 to 0.959 on all 269 corpus rows. On 238 unseen prompts, the repaired 1B passes at 0.536 vs the 600M's 0.428 (p = 0.004). Embedding-drift analysis shows 97.7% of the bf16 embedding table remains bit-identical after repair, meaning the fix lives in the network's internal representations, not in the token embeddings themselves. Three orders of magnitude fewer tokens than the failed 6B-token phase. Integrity deserves close reading. The sample sizes are small: 6 verbatim examples for the initial diagnostic, 269 corpus rows, 238 unseen prompts. The factorial analysis of repair configurations confirms all install the format, but the authors are honest that suppression benefits (the model not over-triggering on negative prompts) show seed sensitivity and remain a hypothesis. Both models over-trigger — the 600M answers negative prompts without a call only 9% of the time, the repaired 1B only 17%. This is a known limitation, not hidden. The paper's real contribution is not the repair but the diagnostic ladder itself: verbatim reproduction, first-token probing, embedding drift, negative-prompt suppression checks. Each costs minutes of CPU time. The argument is that any tool-use claim on a small model should be gated by these checks before anyone trusts a BLEU score. The ladder is method-agnostic and architecture-agnostic — it works wherever special tokens gate structured output. What's missing is scale. This is demonstrated on two models in one language (Spanish) in one domain (security). The diagnostic ladder is plausible as a general tool, but generalization to other languages, domains, and model families is untested. The authors acknowledge this. The factorial analysis, while rigorous in structure, operates on small cells with seed sensitivity. This is a proof-of-concept with a strong central finding, not a large-scale empirical study.