Imagine a courtroom where the witness is never allowed to say how confident they feel. Instead, the judge watches the evidence trail — how many corroborating documents were found, whether the citations actually support the claims, how many independent passes turned up the same answer — and assigns confidence from that paper trail. That is what BITEM's pipeline does for question answering: the model answers, but an external orchestrator decides how much to trust it. The committed claim is architectural, not accuracy-based: you can build a better-calibrated QA system by computing confidence from observable pipeline signals (retrieval depth, entailment checks, multi-pass agreement) than by asking the model to rate itself. The system runs each question three or four times, each pass retrieving from a corpus stripped of previously seen passages, building an evidence dossier the orchestrator evaluates through an entailment cascade. A claim is admitted only once checked against its cited passage. Confidence is never a model self-report. On the NTCIR-19 R2C2 shared task, BITEM's retrieval runs placed 4th and 5th of 22. Pooling multiple passes gained 0.0709 nDCG@20, with the largest lift on multi-hop and post-processing-heavy questions, where their pooled run ranked first in the field. Twelve runs from other teams were built on passages BITEM's retrieval supplied — a quiet infrastructure contribution. The raw pipeline achieved 92.19% accuracy (6th of 25) and an HMR of 0.4915 (13th). But simple hand-crafted rules over the same pipeline signals — no additional model calls, no additional retrieval — pushed accuracy to 93.75% (5th) and HMR to 0.6985 (9th). The delta is stark: the information to calibrate confidence was already sitting in the pipeline's own exhaust. The integrity picture is solid for a shared-task participant paper. NTCIR-19 R2C2 is a community-organized evaluation with organizer-controlled topics and pooled relevance judgments — not a benchmark the authors chose after seeing results. Rankings are public and comparative across 22-25 systems. The weakness is that this is a single evaluation event on a movie corpus; generalization to other domains is untested. The authors are transparent about where they rank and where they don't. BITEM also proposes accHMR (accuracy × HMR) as an alternative metric, arguing that HMR alone can reward a system for being wrong with low confidence. On accHMR, their revised rules score 0.6549, 5th of 25. It is a sensible diagnostic proposal, though whether the community adopts it depends on organizer buy-in. The most interesting forward direction is the one the authors name themselves: replacing hand-crafted rules with a learned model over the pipeline's numeric signals. They have the signal space mapped out — retrieval overlap across passes, entailment cascade outcomes, evidence depth — but fitting a model requires more labeled data than a single shared-task cycle provides. This is almost certainly a compute-and-data limitation, not a theoretical gap. The obvious next experiment is a lightweight classifier (logistic regression, small gradient-boosted tree) trained on pipeline features to predict answer correctness, tested on a held-out question set. The broader argument this paper participates in is whether LLM confidence should come from the model itself (logit-based, verbalized) or from external signals in the pipeline that produced the answer. BITEM's results are a concrete data point for the external-signal camp: pipeline exhaust alone, processed by simple rules, produced a calibration jump from 13th to 9th on HMR with no additional inference cost. That is a practical result, not a theoretical one.