TokenRouter is an efficient and developer-friendly serving system for token-level LLM routing. You describe routing logic through three interfaces: route(), send(), and receive(), TokenRouter efficiently serves it for you.
Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 2.01β64.15Γ higher decoding throughput than existing systems.
Feel free to star the repo or cite the paper if you find it interesting.
@article{fu2026tokenrouter,
title={TokenRouter: Efficient Serving System for Token-Level LLM Routing},
author={Tianyu Fu and Tengxuan Liu and Ruoxi Wang and Yixin Dong and Yi Ge and Yichen You and Yu Wang},
journal={arXiv preprint arXiv:2610.12242},
year={2026},
}[2026/10] π Our paper is now available on arXiv. Explore the project page for interactive explanations and results.
[2026/09] π Our paper is accepted by NeurIPS 2026!
conda create -n tokenrouter python=3.10
conda activate tokenrouter
pip install -e .TokenRouter is configured through a single YAML file. Example configs are provided under config/, and runnable usage examples are provided under examples/.
Launch a TokenRouter server with a routing config:
python -m tokenrouter.launch_server --config_path config/r2r/Qwen3-0.6B_Qwen3-32B.yamlTo place one model on each node, keep both model definitions in the shared
config, add system.peer_endpoints, and select the model launched by each
process. Replace <QUICK_NODE_IP> and <REFERENCE_NODE_IP> with the reachable IP addresses of the corresponding nodes
system:
entry_model: quick
peer_endpoints:
quick:
bind: "tcp://0.0.0.0:30001"
advertise: "tcp://<QUICK_NODE_IP>:30001"
reference:
bind: "tcp://0.0.0.0:30001"
advertise: "tcp://<REFERENCE_NODE_IP>:30001"# On quick node
python -m tokenrouter.launch_server --config_path config.yaml --local_model quick
# On reference node
python -m tokenrouter.launch_server --config_path config.yaml --local_model referenceOnce the server is running, send a chat completion request with either the OpenAI client example or a plain requests example:
python examples/openai_chat_completion.py
python examples/requests_chat_completion.pyUse TokenRoutingEngine when you want to run TokenRouter directly from Python without starting a separate HTTP server:
from tokenrouter.entrypoints.engine import TokenRoutingEngine
with TokenRoutingEngine("config/r2r/Qwen3-0.6B_Qwen3-32B.yaml") as engine:
response = engine.generate(
prompt="<|im_start|>user\nSay hello in one short sentence.<|im_end|>\n<|im_start|>assistant\n",
sampling_params={"temperature": 0, "max_new_tokens": 256},
)
print(response["text"])See examples/engine_generate.py for a complete example:
python examples/engine_generate.pySelect one config and one workload before launching the server.
| Method | Config path |
|---|---|
| R2R | config/r2r/Qwen3-0.6B_Qwen3-32B.yaml |
| CITER | config/citer/Qwen3-0.6B_Qwen3-32B.yaml |
| Co-LLM | config/collm/Qwen3-0.6B_Qwen3-32B.yaml |
| R-Stitch | config/rstitch/Qwen3-0.6B_Qwen3-32B.yaml |
| ME | config/me/Qwen3-0.6B_Qwen3-8B_Qwen3-32B.yaml |
| Workload | Dataset path |
|---|---|
| AIME | evaluation/datasets/aime.json |
| AIME 8K output budget | evaluation/datasets/aime_8k.json |
| SWE-smith multi-round trajectories | evaluation/datasets/swesmith_traj_multi_round.json |
If multiple models share the same physical GPU, CUDA MPS is recommended for reproducing overlapping serving behavior and improving concurrent scheduling performance.
mkdir -p tmp/mps_${METHOD}_${PORT}
export CUDA_MPS_PIPE_DIRECTORY=$PWD/tmp/mps_${METHOD}_${PORT}/pipe
export CUDA_MPS_LOG_DIRECTORY=$PWD/tmp/mps_${METHOD}_${PORT}/log
nvidia-cuda-mps-control -dStart the TokenRouter server with the selected routing config.
python -m tokenrouter.launch_server --config_path $CONFIG_PATHRun the workload on the server and write per-request measurements to the output directory.
bash evaluation/test_throughput.sh \
$CONCURRENCY \
$PORT \
$NUM_INPUTS \
$REPEATS_PER_INPUT \
$MAX_TOKENS \
$OUTPUT_DIR \
$DATASET_PATHThe results will be written to ${OUTPUT_DIR}/result.csv.
Use evaluation/report.py to summarize an existing result CSV without rerunning the benchmark.
python evaluation/report.py --result ${OUTPUT_DIR}/result.csv| Method | Scheduler Classes |
|---|---|
| R2R | r2r_slm + r2r_llm |
| CITER | citer_slm + citer_llm |
| Co-LLM | collm_base + collm_assist |
| R-Stitch | rstitch_slm + rstitch_llm |
| GlimpRouter | glimprouter_slm + glimprouter_llm |
| ME | random_routing on small + medium + large |
| Query-level routing | query_level_routing |
| Random baseline | random_routing |
Code structure and instructions for adding custom routing methods.
Expand to see
tokenrouter/: The main serving package.entrypoints/: HTTP serving, OpenAI-compatible chat completions, and the Python engine.managers/: Scheduling, request handoff, tokenization, and detokenization.managers/schedulers/: Routing schedulers and the scheduler registry.algorithms/: Method-specific routing logic and helpers.model_executor/: Model execution and CUDA graph support.
config/: YAML configurations for routing methods and supported model-level arguments.examples/: Server clients and Python engine examples.evaluation/: Workloads, throughput benchmarks, reporting, and delayed-batching optimization.
TokenRouter represents each cooperating model with its own scheduler class. A new routing method usually adds one scheduler for each model. Every scheduler should subclass BaseTokenRoutingScheduler, defined in tokenrouter/managers/schedulers/base_scheduler.py. Place scheduler classes in tokenrouter/managers/schedulers/ and keep method-specific helper logic under tokenrouter/algorithms/. For each scheduler class, implement three key interfaces: route(), send(), and receive().
route(batch, result) is called after each forward pass and returns one decision per request: a list of model names or a 1-D torch.Tensor with shape (batch.batch_size(),). Each element corresponds to the request at the same index in batch.reqs.
Return one routing decision per request:
- Return a model name to route explicitly, for example
"small","medium", or"large". Returning the current model name continues decoding locally. - For existing two-model schedulers,
True/1still delegates to the only peer andFalse/0still decodes locally.
send(req, destination) builds the PeerReq message sent to an addressed
scheduler. Use PeerReq.construct_from_req(...) to convert the local Req
into a peer message, passing the source scheduler name, destination scheduler
name, and the token ids that should be transferred:
current_token_ids = req.origin_input_ids + req.output_ids[:-1]
req_to_send = PeerReq.construct_from_req(
req=req,
source=self.server_args.server_name,
destination=destination,
current_token_ids=current_token_ids,
entry_model=self.server_args.entry_model_name,
handoff_seq=req.handoff_seq,
)The default send() implementation uses req.origin_input_ids + req.output_ids as current_token_ids, so the peer receives the prompt plus all generated tokens currently owned by this scheduler.
receive(msgs) processes PeerReq messages from the peer scheduler. By default, receive() accepts all tokens sent by the peer. Override receive() when the method needs custom preprocessing before applying peer tokens.
Register the scheduler in tokenrouter/managers/schedulers/init.py by adding it to SCHEDULER_REGISTRY. The registry key becomes the YAML config value for scheduler_class.
The delayed-batching threshold can be optimized using evaluation/calculate_delayed_batching.py. For example, to optimize the delayed-batching threshold of R2R on Qwen3-0.6B and Qwen3-32B, run the following command:
python evaluation/calculate_delayed_batching.py \
--M 2 --N $CONCURRENCY --latencies 6.0 27.9 \
--P '[[0.35, 0.65], [1, 0]]' \
--C '[[1, 0], [1, 0]]' \
--delta 0.1Each cooperating model is configured independently inside the YAML config. TokenRouter extends sglang.srt.server_args.ServerArgs with additional routing-specific fields defined in tokenrouter/server_args.py. All the supported model-level configuration keys are listed in config/supported_server_args.yaml.
Explore more efficient LLM projects from us:
|
C2C
Direct communication through KV-caches |
R2R
Token-level routing for reasoning LLMs |
TaH
Selective latent thinking for reasoning LLMs |
FrF
Efficient video token reduction for LVLMs |
MoA
Mixture of sparse attention for LLMs |
