SaveRouter trains LLM routers on about 33 to 41% of the feedback

LLM routing saves money at serving time by sending each query to a suitable model while keeping response quality. The paper starts from a cost that is easy to overlook: to learn a router, teams often have to run several candidate models on historical queries and record how well each one answered. That feedback collection is a real expense paid before deployment. The authors say existing work mostly measures serving-time efficiency and does not ask whether the resulting savings are big enough to recover that upfront spend.
They also report an observation: routing quality often saturates well before all query-model feedback has been collected. In their reading, collecting feedback densely can be economically over-provisioned.
Their answer is SaveRouter, a sparse-supervision routing framework. It does three things. It selectively acquires the model feedback that is informative. It shares capability information across related queries. And it keeps query-level refinement so that routing stays fine-grained. The evaluation counts supervision expenditure and later serving-time savings together, rather than looking at serving savings alone.
Across four routing benchmarks, the main setting uses only about 33 to 41% of the available training feedback while maintaining competitive or better routing quality. It also reduces the break-even deployment volume (the amount of serving needed before savings repay the training cost) by approximately 1.9 to 9.5 times compared with the fastest conventional router.
A further analysis adds a caveat to the "more data is better" intuition. Acquiring more supervision is not always economically preferable: the supervision level that minimizes serving cost can differ from the one that reaches the earliest payback. The code is publicly available in the LAMDA-Model-Reuse GitHub organization, at github.com/LAMDA-Model-Reuse/SaveRouter.
Key facts
- SaveRouter is a sparse-supervision routing framework that selectively collects informative model feedback and shares capability information across related queries, keeping query-level refinement.
- In the main setting it uses only about 33 to 41% of the available training feedback across four routing benchmarks, with competitive or better routing quality.
- It reduces break-even deployment volume by approximately 1.9 to 9.5 times compared with the fastest conventional router.
- The paper evaluates routing by counting supervision cost and serving-time savings together, not serving savings alone.
- More supervision is not always economically better: the level that minimizes serving cost can differ from the level with the earliest payback. Code is public on GitHub.
Why it matters
Routing is sold on serving savings, but a router has to be trained first, and training can mean running several candidate models over historical queries. This paper puts that upfront cost into the accounting and asks when the savings actually pay it back. Break-even deployment volume is the yardstick it uses. The reported result is that a sparse approach can reach the payback point with roughly 1.9 to 9.5 times less deployment volume than the fastest conventional router.
Who it affects
Mainly teams that serve several LLMs and are deciding whether to build a router to cut serving cost, and researchers working on routing and model selection. The argument is most relevant to anyone for whom collecting query-model quality feedback is a meaningful expense before launch.
How to use it
The authors publish code at https://github.com/LAMDA-Model-Reuse/SaveRouter. The practical idea is to avoid collecting full feedback for every query against every candidate model: acquire the informative feedback selectively, share capability information across related queries, and keep query-level refinement. The paper also suggests checking payback, not only serving cost, when deciding how much supervision to buy.
How solid is it
The results come from the authors' own evaluation across four routing benchmarks. The figures are ranges: about 33 to 41% of the available training feedback, and approximately 1.9 to 9.5 times lower break-even volume. The source does not name the four benchmarks or the candidate models, and does not say which value belongs to which benchmark. It also gives no numeric quality metric; quality is described only as competitive or better. The baseline is called the fastest conventional router and is not named.
Risks and caveats
No absolute cost figures such as dollars or tokens are given, so the real-world savings for a given deployment cannot be read off the paper. Quality is described only as competitive or better, with no accuracy number. The 33 to 41% figure applies to the main setting. The authors themselves note that more supervision is not always economically preferable, and that the level minimizing serving cost can differ from the one with the earliest payback, so the best setting depends on what a team is optimizing.
“Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure.”
— SaveRouter paper abstract