⚡TokenRouter Serves Token-Level LLM Routing Up to 64x Faster
TL;DR
Tsinghua's TokenRouter is a serving system for routing individual tokens between models in one generation. It reports 2.01x to 64.15x higher decoding throughput than existing systems across routing algorithms and model pairs.
Tsinghua's TokenRouter is a serving system for routing individual tokens between models in one generation. It reports 2.01x to 64.15x higher decoding throughput than existing systems across routing algorithms and model pairs.
Key Points
Targets step desynchronization and batch-admission delays that single-model servers hit under token-level routing
Developers write routing logic per request; the runtime launches one subserver per model
Delayed-batching scheduler with hyperparameters set by a throughput model
Code released at github.com/thu-nics/TokenRouter
Posted to arXiv Oct. 8 by the Tsinghua NICS-EFC group
Why It Matters
Mixing a small and a large model token by token only pays off if the serving stack keeps GPUs busy. This is the missing infra piece, and the code is open.
Quick Facts
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.