Language model routing improves by using semantic probabilities

SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing

Artificial IntelligenceComputation and LanguageMachine Learning

Summary

Choosing the right language model for a question is tricky because most methods don't clearly explain what the question needs. The authors introduce SeLMRoute, which first breaks down questions into simple, understandable parts and estimates these parts as probabilities. Then, a light model guesses how well different large language models might answer based on these parts. This approach helps pick better language models for questions and works well across many tests.

What this means in practice

  • For ai platform developers: Select the best large language model for each user query by building routers that use interpretable semantic indicators rather than just embeddings.
  • For cloud service operators: Optimize cost and performance trade-offs when deploying multiple language models by integrating semantic evidence into routing decisions.

Authors

Vasilis Perifanis, Nikolaos Pavlidis, Symeon Symeonidis

Abstract

Large language model (LLM) routing aims to select the most suitable model for each incoming query. Most existing routers learn this decision directly from query embeddings, model representations, preference data, or clusters of similar examples. Such approaches can be effective, yet the representation used for routing rarely states what a query actually requires. We introduce SeLMRoute, a routing framework that separates the extraction of candidate-independent semantic evidence from the learning of candidate performance and the application of deployment objectives. A decision model first evaluates a set of interpretable questions about the query, such as its reasoning requirements and use of external knowledge, with each judgment retained as a probability distribution. The resulting probabilistic semantic state is used by a lightweight supervised router to estimate candidate model performance. Routing objectives are applied after performance estimation, which allows the same semantic state to support performance-oriented and cost-aware decisions. On the LLMRouterBench (15 datasets, 20 candidate models, 11,481 queries), SeLMRoute achieves an average accuracy of $72.08\% \pm 0.45$, while grouped five-fold out-of-fold evaluation reaches $72.64\%$, compared with $69.23\%$ for the strongest fixed candidate. The representation achieves the highest mean performance among the evaluated semantic, dense, lexical, and domain-level representations. In a separate 13-model performance-cost setting, SeLMRoute improves performance in all five grouped splits, with a mean PerfGain of $2.66\%$. Our code is available at https://github.com/Indigma-Innovations/SeLMRoute.