Imagine you hired a contractor to renovate your kitchen. They tell you it's done. You could ask them "how confident are you it's right?" — but that's one number from the person who did the work. A smarter move: walk through a checklist. Were permits pulled? Is the plumbing to code? Do the outlets work? Each sub-check has its own evidence, and the overall confidence emerges from combining them. That's the core mechanism of Confidence Reasoning Graphs. The committed claim: CRGs estimate the probability that an LLM agent accomplished its task from a single trajectory — no model internals, no training data, no repeated rollouts — by decomposing the top-level success claim into a directed acyclic graph of sub-claims grounded in trajectory evidence, scoring each leaf, and aggregating upward. This is genuinely new as a structured, inference-time-only framework. Prior approaches either ask the model "how sure are you?" (verbalized confidence), run the agent multiple times and count successes (sampling-based), or train a surrogate classifier on internal activations (white-box). CRGs sidestep all three limitations. The results are tested across three agentic benchmarks (the paper mentions three though specific names aren't given in the abstract), three backbone LLMs, and three agent frameworks — a 3×3×3 grid that gives real coverage. CRGs produce better-calibrated confidence and stronger risk-aware decision-making than all baselines. The most interesting finding isn't the headline win; it's the diagnostic: a white-box surrogate baseline appeared well-calibrated by Expected Calibration Error but provided near-chance discrimination. That's a sharp methodological contribution — calibration error alone is a misleading metric for this problem, because a model can be well-calibrated on average while being unable to distinguish success from failure on individual instances. Architecturally, CRGs belong to the inference-time reasoning family — no gradient updates, no finetuning, no privileged access to logits or hidden states. The method leans on the LLM's own reasoning capability to construct and evaluate the graph. This makes it model-agnostic and deployable with API-only frontier models, which is the practical setting most teams actually face. The decomposition step is reminiscent of chain-of-thought prompting but structured as a graph rather than a linear chain, with explicit aggregation rather than implicit coherence. The ablation story is clean: improvements come from claim-level confidence estimation and the aggregation function, not from graph construction alone. This means the graph structure is load-bearing for auditability but the scoring and combining of leaf claims is what drives calibration improvements. That's a useful separation — it tells you what to invest engineering effort in if you're building on this. The integrity picture is reasonable for a methods paper but has gaps. The 3×3×3 evaluation matrix is strong coverage, but benchmark names aren't specified in the abstract, there's no mention of pre-registration, and the paper is 34 pages with 11 tables — suggesting thorough internal validation but no independent replication yet. The honest limitation is that CRGs add inference-time compute (LLM calls to decompose, score, and aggregate) and the cost-benefit tradeoff versus simpler baselines at scale isn't characterized. The practical upshot: if you're deploying LLM agents in consequential domains — customer support, code generation, document processing — and you need a confidence signal to decide when to route to a human, CRGs give you something that existing approaches don't: a structured, auditable decomposition of why the agent might have failed, not just a number. The claim graph is the artifact a human reviewer can actually inspect at decision time.