Imagine you're a doctor reviewing lab results. You notice two tests contradict each other — one says the patient's iron is dangerously low, the other says it's normal. A good doctor flags the discrepancy, orders a retest, and tells the patient there's uncertainty. A bad doctor picks one result, writes a confident diagnosis, and moves on. This paper catches LLM agents doing the bad-doctor thing, systematically. The committed claim: task accuracy and epistemic humility are decoupled in LLM agents. The authors introduce an evaluation framework — ISE (Identify, Solve, Escalate) — that tracks not just whether an agent gets the right answer, but whether it recognizes contradictions between its parametric knowledge and retrieved evidence, whether it resolves those contradictions, and whether it communicates residual uncertainty to the user. Across four agents tested in two conflict settings, high accuracy configurations routinely detected conflicts in their intermediate reasoning steps but then suppressed that uncertainty in their final outputs. The finding is clean and specific: the failure is not in perception but in communication. The evaluation design is more careful than most agent benchmarks. The authors construct two conflict settings: controlled conflicts where they deliberately inject contradictory evidence into retrieval results, and naturally occurring conflicts that emerge during multi-step agentic execution on real tasks. Crucially, each conflict condition is paired with a matched no-conflict control, so the behavioral delta is attributable to the conflict itself rather than task difficulty. Four agents are evaluated, though the paper is light on naming specific backbone models in the abstract. The trajectory-level analysis is where this gets interesting. Rather than just scoring final answers, the authors trace the agent's reasoning chain step by step. They find a striking pattern: agents frequently identify conflicts early in execution — step 2, step 3 — but the conflict signal degrades through subsequent steps. By the time the agent produces its final answer, the uncertainty has evaporated. This is the reasoning-chain equivalent of organizational memory loss: the information was there, it just didn't survive the pipeline. The model-level intervention results reveal a genuine tradeoff, not a free lunch. When the authors apply interventions to improve epistemic humility — likely prompting strategies or calibration techniques — they find that EH improves but task accuracy drops. This is the paper's most important finding for practitioners: you cannot simply instruct an agent to be more humble without accepting a performance cost. Epistemic humility is not a property of the model alone but emerges from the interaction among the backbone model, the agent harness, and the evaluation environment. The integrity picture is mixed but honest. The controlled-conflict setup is rigorous — matched controls, trajectory-level measurement, behavioral decomposition into three distinct dimensions. But the validation is fundamentally same-team: the authors define the ISE framework, build the evaluation, and grade the results. There's no independent replication, and the ISE operationalization itself is a judgment call that could be contested. The paper is accepted at EMNLP 2026, which provides peer-review filtering. For anyone building agentic systems that touch high-stakes decisions — medical, legal, financial — this paper names a failure mode that existing accuracy benchmarks completely miss. The practical implication is stark: an agent that scores 90% on your task benchmark may be confidently wrong on the other 10%, even when its own reasoning chain contains the evidence that something is off. The ISE framework gives you a vocabulary and measurement approach for auditing this specific blind spot.