Wire flash
Tsinghua paper: TokenRouter serving system boosts LLM throughput up to 64.15x
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
A new paper from Tsinghua University introduces TokenRouter, a serving system that performs per-token routing between small and large language models, achieving up to 64.15 times the throughput of existing setups. Current popular serving frameworks such as vLLM and SGLang run one model per request, causing every step to wait for the slower model when two models share an answer. TokenRouter gives each model its own server and allows them to pass work back and forth, handing a half-written answer between models while preserving the model's memory of the text so far (the KV cache). It also holds requests momentarily so each model works on larger batches. Across five routing methods, throughput increased between 2.01 and 64.15 times compared to the stronger existing setup. The paper is available on arXiv under the title 'TokenRouter: Efficient Serving System for Token-Level LLM Routing'.
Source report
A new paper from Tsinghua University introduces TokenRouter, a serving system that enables per-token routing between small and large language models, achieving up to 64.15× the throughput of existing setups.
The Problem with Current Systems
Popular serving frameworks such as vLLM and SGLang run one model per request. When two models share an answer, each step must wait for the slower model, creating bottlenecks.
How TokenRouter Works
TokenRouter assigns each model its own server and allows them to pass work back and forth. Key features include:
- Hands off a half-written answer between models while preserving the model's memory of the text so far (the KV cache)
- Holds requests momentarily so each model can work on larger batches
Performance Results
Across five routing methods, TokenRouter achieved throughput improvements ranging from 2.01× to 64.15× compared to the strongest existing setup.
Paper: "TokenRouter: Efficient Serving System for Token-Level LLM Routing" Source: arxiv.org/abs/2610.12242
Source
rohanpaul_aiNeutral / independent