Skip to content
daily-hour-news·

⚡TokenRouter Serves Token-Level LLM Routing Up to 64x Faster

TL;DR

Tsinghua's TokenRouter is a serving system for routing individual tokens between models in one generation. It reports 2.01x to 64.15x higher decoding throughput than existing systems across routing algorithms and model pairs.

Tsinghua's TokenRouter is a serving system for routing individual tokens between models in one generation. It reports 2.01x to 64.15x higher decoding throughput than existing systems across routing algorithms and model pairs.

TokenRouter Serves Token-Level LLM Routing Up to 64x Faster — daily-hour-news

Key Points

1

Targets step desynchronization and batch-admission delays that single-model servers hit under token-level routing

2

Developers write routing logic per request; the runtime launches one subserver per model

3

Delayed-batching scheduler with hyperparameters set by a throughput model

4

Code released at github.com/thu-nics/TokenRouter

5

Posted to arXiv Oct. 8 by the Tsinghua NICS-EFC group

Why It Matters

Mixing a small and a large model token by token only pays off if the serving stack keeps GPUs busy. This is the missing infra piece, and the code is open.

Quick Facts

TokenRouterLLM servinginferenceroutingTsinghuathroughputopen source

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Also get