Imagine you hire a junior analyst and hand them a spreadsheet. They run a pivot table, get a number, and put it in the report. The number is wrong — a corrupted cell, a misjoined table — but the formula executed without errors. No red text, no stack trace. The analyst trusts the output because the tool said it worked. That is exactly what happens inside tool-augmented LLM agents, and this paper builds a systematic way to measure how often and how badly. The committed claim: current tool-augmented data agents exhibit "blind compliance" — they accept plausible but incorrect tool outputs at alarming rates, and existing retry mechanisms only partially fix this. ToxicBench is the first benchmark explicitly designed to separate two failure modes that prior work conflated: whether the agent bothered to CHECK the evidence, and whether it ultimately ADOPTED the correct answer. That distinction turns out to matter a lot. An agent can check and still adopt the wrong answer, which is a worse failure mode than never checking at all because it looks like diligence. The benchmark pairs clean and poisoned observations across four error types — numerical errors, label errors, schema errors, and retrieval errors — over fixed source data, so you can isolate the effect of the poison from the difficulty of the task. Across 118 tasks evaluated with GPT and three adapter configurations (Base, CoT, ReAct), poisoning drops task success by 26 to 39 percentage points depending on adapter. The sharpest finding: ordinary retries help when poisoning is one-shot, but under repeated poisoning, agents adopt wrong answers even after performing checking steps. The retry mechanism creates an illusion of robustness. Architecturally, this is a benchmark paper, not a new model. The agents under test are GPT-based with standard tool-calling adapters. The poisoning mechanism operates at the tool-output layer — the tool executes correctly in terms of code, but returns corrupted evidence. This is a realistic threat model: real-world APIs return stale data, databases have integrity issues, and scraping pipelines silently break. The design choice to hold source data fixed while toggling clean/poisoned observations is the methodological backbone — it makes the comparison controlled rather than confounded. Integrity is decent for a benchmark paper. Controls on three public tables isolate evidence effects. The scorer was frozen before comparison with human annotations on 200 trajectories, finding 96% task-success agreement. That is a strong signal that the automated evaluation isn't gaming itself. Human judgments confirm the retry gains over Base and the troubling wrong-answer adoption. Code and benchmark materials are released. The main integrity gap: all evaluation is on GPT models via one provider. We don't know if Claude, Gemini, or open-weight models show the same blind compliance patterns or if this is partly a GPT-specific behavior. The milestone question is where this gets interesting for practitioners. Right now, the paper demonstrates the problem and quantifies it. The next concrete number to watch is whether an intervention — verification chains, self-consistency checks, or grounded tool-output validation — can recover more than half of the 26-39 point drop without doubling inference cost. That would convert this from a diagnostic into a defense. The authors gesture toward this but don't run it, which reads as saving it for the next paper rather than a negative result they're hiding. The bigger field fight here is whether tool-augmented agents need explicit verification architectures or whether scaling and better prompting will naturally reduce blind compliance. This paper lands firmly on the side that says scaling alone won't fix it — the failure mode is structural, not capacity-limited. An agent that executes code successfully has no built-in reason to doubt the output, regardless of how capable the underlying LLM is. That argument, if it holds across models and scales, has real implications for how production agent systems are designed.