You know that feeling when you start a home renovation and the contractor says "about $20K" but the final invoice lands somewhere between $15K and $180K depending on what they find behind the walls? That's the current state of LLM agent cost prediction. An agent tackling the same coding task can consume wildly different token counts across runs because each tool call's output reshapes the growing context window, and every subsequent call re-reads that ballooning context. TokenCast treats this problem the way a seasoned project manager treats construction: not by predicting the whole job upfront, but by learning the cost signature of each phase and how each phase inflates the cost of every phase after it. The core mechanism is a composable cost representation. Each execution segment — a single tool call or reasoning step — gets a learned embedding that encodes two things: its own token consumption and the context growth it introduces for downstream segments. When you compose adjacent segments, the model captures the multiplicative effect of context accumulation: early verbose tool outputs don't just cost tokens themselves, they inflate the input size of every later LLM call. This is the key structural insight. Token cost in agentic execution is not additive — it's compounding, because context is re-read at every step. TokenCast works in two phases. Offline, it trains on execution traces to learn segment-level cost representations. Online, as a run unfolds, newly observed segments refresh the forecast without making any additional LLM calls. The inference overhead is trivial: 32.8 milliseconds mean cumulative prediction time per run on SWE-bench Verified. You're essentially getting a continuously updated cost estimate for free relative to the cost of the agent itself. The evaluation is broad by agent-cost-prediction standards. The authors test across 4 task suites and 6 agent models, yielding 96 evaluated combinations. Against the strongest comparator (not named in the abstract but presumably a trajectory-level regression baseline), TokenCast achieves a mean absolute error reduction of 14.5%. The practical payoff shows up in budget control: in offline replay experiments, TokenCast uses 21.3% fewer tokens than a fixed-budget policy while completing the same traces. That's real money at scale — if you're running thousands of agent tasks daily, a 21% reduction in token spend is a significant operational savings. The honest limitation is that this is still replay-based validation for the budget-control claim. The paper demonstrates that if you had known TokenCast's forecasts during execution, you could have made better stopping decisions. The next step — actually deploying TokenCast as a live controller that terminates or redirects agent runs in real time — is the experiment that would close the loop from forecasting to intervention. What makes this paper worth attention is not that it solves a glamorous problem but that it names a real operational pain point that will only intensify. As agents get more autonomous and context windows grow, the variance in token consumption per task will increase, not decrease. Someone needs to build the metering infrastructure for agentic compute. TokenCast is an early serious attempt at that metering layer, and the composable-segment approach is the right structural bet because it respects the compounding nature of context costs.