Imagine you're a postal sorting office that handles letters for two cities. The old system puts every letter on the same conveyor belt, and when half the letters need City A and half need City B, you stop the whole belt, reroute, restart. That's how current LLM serving systems handle token-level routing — they were built assuming one model, so splitting tokens across two models creates stop-start chaos. TokenRouter redesigns the conveyor: each city gets its own belt, a smart dispatcher feeds letters to whichever belt has capacity, and a mathematical model tells you the optimal moment to seal each batch. The committed claim: token-level LLM routing (where individual generated tokens can be assigned to different models mid-sequence) is algorithmically promising but practically unservable on existing infrastructure. TokenRouter is the first serving system designed specifically for this workload, and it achieves 2.01–64.15× higher decoding throughput than adapted single-model systems. That range is wide because baseline systems choke differently under different routing patterns — the worst cases are where step desynchronization between models is most severe. The core architecture insight is the separation of programming model from execution model. Developers write routing logic as if they're handling a single request ('request-centric programming'), while the runtime decomposes this into per-model subservers that execute asynchronously ('model-centric execution'). Each subserver runs a delayed-batching scheduler — it waits an optimal duration before sealing a batch, balancing the tension between batching efficiency (bigger batches = better GPU utilization) and latency (waiting too long starves the pipeline). The optimal delay parameters come from a closed-form throughput model, not heuristic tuning. The step desynchronization problem deserves unpacking because it's the actual bottleneck this paper attacks. In standard continuous batching (the vLLM/TGI paradigm), all requests in a batch step through the same model in lockstep. Token-level routing breaks this assumption: request A might need Model 1 for tokens 1-5 and Model 2 for tokens 6-10, while request B does the opposite. Naive implementations either serialize these or pad batches with idle slots. TokenRouter's delayed-batching scheduler accumulates enough same-model work to form efficient batches, with the mathematical model ensuring you're not waiting so long that end-to-end throughput drops. The evaluation spans diverse routing algorithms, workloads, and model pairs — though the abstract doesn't name specific baselines or model sizes. The 2.01× floor suggests that even in easy cases (low routing frequency, similar-sized models), the system design pays off. The 64.15× ceiling likely reflects pathological cases for naive systems where desynchronization is extreme. NeurIPS 2026 acceptance provides peer-review validation, and code is publicly available. The field fight this enters is practical, not theoretical: the ML systems community has been scaling single-model serving (vLLM, TensorRT-LLM, SGLang) while the algorithms community has been developing multi-model routing strategies (RouteLLM, hybrid MoE, speculative decoding). The gap between these two communities is the serving infrastructure for routing. TokenRouter is a bridge paper — it takes algorithmic routing ideas and makes them deployable. What's missing is the cost-quality Pareto analysis on real workloads. The throughput numbers are compelling, but the downstream question is: does token-level routing with TokenRouter actually beat a single larger model on quality-per-dollar? The system enables the experiment, but the paper (at least from the abstract) focuses on the systems contribution rather than the end-to-end economic argument.