You know that feeling when you taste-test five brands of coffee in your kitchen, pick a favorite, then wonder if you'd pick the same one tomorrow? Now imagine you're required to prove your pick would hold up 95% of the time — and you have to prove it for every possible number of brands you might taste, simultaneously. That's the certification problem for test-time scaling curves, and it's far harder than drawing the curves themselves. The core claim: certifying the full accuracy-vs-budget scaling curve for best-of-k sampling requires a simultaneous confidence band covering all budgets at once, and the naive approach (exact binomial at each budget) is absurdly wasteful. On a 100-question benchmark with 64 budgets, that naive design needs 192,000 generated answers to certify accuracy within ±1/32 at 95% confidence. The paper derives the minimax cost of doing this properly and builds an audit that slashes it. The key insight is a variance decomposition. At budget k=64, roughly three-quarters of the variance in whether a selected answer is correct comes from between-question differences — some questions are just harder. A benchmark is a fixed list of questions, so an audit that revisits every question can eliminate that between-question variance entirely. What remains is within-question noise: the randomness of which answers get sampled for a given question. The minimax cost decomposes into three pieces — calibrating score-distribution tails, discriminating between questions, and summing within-question noise along the curve — and the last piece sharpens to variance-of-influence under optimal allocation. The practical tool is a paired audit built on an exponential inequality for two independent draws at the same question. It requires no pilot study, no distributional assumptions beyond independence within a question. On 185 held-out score pools it consumed 0.74× the answers of the cheapest competing certified audit at 64 budgets and 0.53× at 1,024 budgets. On a newly generated MMLU-Pro study, it certified the full curve with 79,133 answers — within 0.6% of a cost-law prediction fitted beforehand from within-question variance. This matters because test-time scaling is the current frontier of LLM capability claims. When labs report 'at k=64 our model achieves X% on MMLU,' the statistical foundation under that claim is usually a single point estimate with no simultaneous coverage guarantee. Anyone selecting a compute budget by reading off the curve is doing post-hoc optimization — the classic multiple-comparisons trap. This paper names the problem precisely and gives it a solution with provable guarantees. The architecture is classical statistics, not ML. The method lives in the family of sequential/adaptive hypothesis testing and minimax estimation theory, drawing on exponential concentration inequalities (the Hoeffding family), stratified sampling, and adaptive allocation. The compute property it exploits is that generating LLM answers is expensive but evaluating the statistical structure of the resulting scores is cheap — so you want to minimize answer generation, not analysis time. The validation is mathematical proof backed by empirical confirmation on real score pools. The 185 held-out pools provide a genuine stress test, and the MMLU-Pro study is a prospective demonstration — the cost law was fitted before the study was run, and the actual cost landed within 0.6% of prediction. That's unusually tight for a statistics paper. The method also extends to pass@k, majority voting, population-level inference, and dependent answers (chain-of-thought), though these extensions are demonstrated more lightly.