TokenRouter claims 2.01-64.15x higher decoding throughput for token-level LLM routing

TokenRouter claims 2.01-64.15x higher decoding throughput for token-level LLM routing

LLM routing spreads inference work across different models, pushing out the cost-quality Pareto frontier of LLM serving. Production systems already use coarse-grained routing at the session or query level. Recent algorithmic work, the authors say, shows that routing at the finer level of individual tokens can bring substantial efficiency and quality gains.

The catch is serving it. According to the paper, current serving systems are built on single-LLM assumptions. Under token-level routing they suffer from severe step desynchronization and frequent batch admission delays, and they also impose high implementation complexity on developers.

To address this, the authors design TokenRouter, described as an efficient and developer-friendly serving system for token-level routed LLM inference. Its guiding principle is "request-centric programming, model-centric execution". Developers describe routing logic from the perspective of a single request. The runtime then launches a separate subserver for each LLM and dispatches requests to them asynchronously. Each subserver runs a delayed-batching scheduler, and the optimal hyperparameters of that scheduler are derived from a mathematical throughput model of the system.

Across diverse routing algorithms, workloads and model pairs, the authors report that TokenRouter reaches 2.01-64.15x higher decoding throughput than existing systems. They say this substantially advances the serving efficiency of token-level LLM routing. The code is available on GitHub at thu-nics/TokenRouter.

Key facts

  • TokenRouter is a serving system for LLM inference in which routing between models happens at the token level, not per session or per query.
  • The authors report 2.01-64.15x higher decoding throughput than existing systems, across diverse routing algorithms, workloads and model pairs.
  • Developers write routing logic from the viewpoint of one request; the runtime launches a subserver per LLM and dispatches requests asynchronously.
  • Each subserver uses a delayed-batching scheduler whose optimal hyperparameters come from a mathematical throughput model.
  • Code is released at github.com/thu-nics/TokenRouter.

Why it matters

Routing queries between models is already common in production, but the paper argues that routing each token is where further efficiency and quality gains lie, according to recent algorithmic work. Serving infrastructure has lagged behind that idea: systems built around a single model run into step desynchronization and batch admission delays when tokens are sent to different models. TokenRouter is an attempt to close that gap on the systems side, and the reported speedup of 2.01-64.15x in decoding throughput is large enough to matter if it holds up.

Who it affects

Teams that serve LLMs and want to mix several models, for example a model pair, inside a single generation. Researchers who design token-level routing algorithms are also affected, since the system is meant to spare developers the high implementation complexity that current serving stacks impose.

How to use it

The code is public at https://github.com/thu-nics/TokenRouter. The programming model is that a developer describes routing logic from the perspective of a single request, and the runtime handles the rest: it launches a subserver for each LLM and dispatches requests asynchronously.

How solid is it

This is a paper with released code, and the headline number is the authors' own measurement: 2.01-64.15x higher decoding throughput than existing systems across diverse routing algorithms, workloads and model pairs. The scheduler's hyperparameters are derived from a mathematical throughput model rather than tuned by hand. The source does not say which setting gives the 2.01x and which the 64.15x figure, nor whether the range is of means, medians or maxima.

Risks and caveats

The gains reported are for decoding throughput only; no latency or output-quality results are reported in the abstract. No specific baseline systems are named, and neither are the routing algorithms, models, model pairs or workloads. No hardware or GPU configuration is given. The width of the range, from about 2x to over 64x, means the real benefit depends heavily on the setup, and independent reproduction would be needed to know what to expect in a given deployment.