Research 01
Decision
Routing, policy, and learning for choosing models, tools, and actions.
Research program
No single model, system, or resource is right for every problem.
These papers develop the principles and mechanisms needed to decide among them—and to test the consequences of those decisions.
Research 01
Routing, policy, and learning for choosing models, tools, and actions.
Research 02
Placement, inference, and resource coordination across heterogeneous infrastructure.
Research 03
Verification, security, and accountability from intent to outcome.
2026
Authors: vLLM Semantic Router Team
Venue: arXiv Technical Report
Signal-driven decision routing for Mixture-of-Modality deployments across cost, privacy, latency, and safety constraints.
PaperAuthors: Huamin Chen, Xunzhuo Liu, Bowei He, Fuyuan Lyu, Yankai Chen, Xue Liu, Yuhan Liu, Junchen Jiang
Venue: arXiv Technical Report
A synthesis of routing, fleet, multimodal, and governance results into the Workload-Router-Pool architecture.
PaperAuthors: Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang, Huamin Chen
Venue: arXiv Technical Report
A security treatment of perception failures in computer-using agents with a dual-channel guardrail for click targets and action reasoning.
PaperAuthors: Huamin Chen, Xunzhuo Liu, Junchen Jiang, Bowei He, Xue Liu
Venue: arXiv Technical Report
OATS improves semantic-router tool ranking under single-digit millisecond CPU budgets without serving-time model inference.
PaperAuthors: Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang, Huamin Chen
Venue: arXiv Technical Report
Adaptive VLM Routing estimates action difficulty and routes each computer-use step to the cheapest model that meets a reliability target.
PaperAuthors: Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang, Huamin Chen
Venue: arXiv Technical Report
Flash Attention, prompt compression, and near-streaming body processing reduce routing latency from seconds to tens of milliseconds.
PaperAuthors: Huamin Chen, Xunzhuo Liu, Yuhan Liu, Junchen Jiang, Bowei He, Xue Liu
Venue: arXiv Technical Report
A queueing-theory-grounded fleet planner and simulator for sizing multi-pool GPU fleets against P99 TTFT targets.
PaperAuthors: Huamin Chen, Xunzhuo Liu, Yuhan Liu, Junchen Jiang, Bowei He, Xue Liu
Venue: arXiv Technical Report
An analytical method for deriving minimum-cost two-pool fleets directly from workload CDFs and P99 TTFT targets.
PaperAuthors: Huamin Chen, Xunzhuo Liu, Yuhan Liu, Junchen Jiang, Bowei He, Xue Liu
Venue: arXiv Technical Report
An analytical result showing context-length routing topology can matter more than pure GPU generation upgrades for tokens-per-watt.
PaperAuthors: Xunzhuo Liu, Hao Wu, Huamin Chen, Bowei He, Xue Liu
Venue: arXiv Technical Report
A framework for conflict detection and prevention when probabilistic ML predicates can silently co-fire in routing policy languages.
PaperAuthors: Huamin Chen, Xunzhuo Liu, Bowei He, Xue Liu
Venue: arXiv Technical Report
A cross-layer extension of the Semantic Router DSL from stateless request routing into multi-step agent workflows.
PaperAuthors: Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang, Huamin Chen
Venue: arXiv Technical Report
Conversational memory and retrieval-grounded routing recover most of a 235B model's performance while cutting effective inference cost by 96%.
PaperAuthors: Xunzhuo Liu, Bowei He, Xue Liu, Haichen Zhang, Huamin Chen
Venue: SIGIR 2026 Industry Track
A real-time verification component for long-document RAG that preserves grounding checks without falling back to truncated validation.
Paper2025
Authors: Chen Wang, Xunzhuo Liu, Yuhan Liu, Yue Zhu, Xiangxi Mo, Junchen Jiang, Huamin Chen
Venue: NeurIPS 2025- MLForSys
A semantic router that classifies queries by reasoning need and selectively applies reasoning only when beneficial.
PaperAuthors: Chen Wang, Xunzhuo Liu, Yue Zhu, Alaa Youssef, Priya Nagpurkar, Huamin Chen
A category-aware semantic caching architecture where similarity thresholds, TTLs, and quotas vary by workload class.
Paper