Imagine you're playing a game of Telephone, but with a twist: the first person deliberately whispers a wrong word. The question isn't whether the message degrades — it will. The question is whether anyone downstream catches the error, corrects it mid-chain, and still arrives at the right answer. PHRBench is a controlled experiment to measure exactly that for large language models, and the answer is sobering: almost nobody catches the error. The committed claim: this is the first benchmark that decomposes post-hallucination reasoning into behaviorally distinct trajectories — not just 'did the model get the right answer despite the wrong premise,' but HOW did the model's reasoning path interact with the hallucinated information? The taxonomy matters. PHRBench categorizes every response into Hallucination Compliance (the model swallows the false premise whole), Hallucination Avoidance (the model sidesteps without explicitly correcting), and Heuristic Correction (the model identifies and overrides the hallucination). Only corrections that also reach the correct final answer count as 'insightful trajectories.' Across 4,820 controlled instances and 18 models spanning four domains, successful recovery is rare — and when it does happen, it correlates with more frequent belief updates within the reasoning chain. The architectural insight here is straightforward but consequential: these are prompt-level perturbation experiments on frozen models. No fine-tuning, no mechanistic probing — just carefully constructed inputs where the ground truth is known, the hallucinated premise is injected, and the resulting chain-of-thought is classified. The four domains provide some breadth, though the abstract doesn't specify them. The evaluation pipeline itself is the contribution: a behavioral taxonomy plus a controlled injection methodology. The most provocative finding is the predictor. A lightweight classifier trained on properties of the hallucinated prompt — not on the model's internal states, not on the reasoning chain, just on the prompt itself — achieves 0.847 AUROC at predicting whether recovery will occur. This implies that recoverability is substantially a property of how the hallucination is presented, not just which model encounters it. If that holds up, it means you could build a pre-flight check on multi-stage LLM pipelines: flag prompts where hallucination propagation is likely to be catastrophic because recovery probability is low. Integrity-wise, this is a well-scoped empirical study with clear limitations. The 18-model spread is healthy for a benchmark paper. The 4,820 instances provide reasonable statistical power. The behavioral taxonomy is defined a priori (the categories exist before the experiments), which is good practice. But the predictor's 0.847 AUROC needs scrutiny: what features drive it? Prompt length, domain, contradiction salience, linguistic markers? The abstract is silent on feature analysis, which is the obvious follow-up. The ladder position is interesting. Prior work on hallucination has focused either on detection (is this output hallucinated?) or on aggregate robustness metrics (does accuracy drop when context is noisy?). PHRBench's contribution is the response-level trajectory decomposition — studying the reasoning path, not just the endpoint. This sits alongside work like TruthfulQA and HaluEval but asks a fundamentally different question: not 'does the model hallucinate' but 'what does the model do when hallucination is already in the room?' The milestone question is where things get practical. For multi-stage agentic systems — retrieval-augmented generation, tool-use chains, multi-agent debates — hallucination propagation is a deployment-blocking risk. The implicit target is a recovery rate high enough to trust LLM chains without human-in-the-loop verification at every stage. We're not there. But the predictor result suggests a viable interim architecture: predict recoverability, and route low-recovery-probability chains to human review or alternative pipelines.