Text features replace neural embeddings for fast and clear routing

Routing Without Embeddings: Fast And Interpretable Routing With Regular Expressions

Machine Learning

Summary

Large language models usually understand questions by turning them into numbers using big, complex tools, but bigger tools didn’t really work better for sorting questions. This paper shows that simpler patterns can be found directly by looking at the text using regular expressions, which are like search rules for text. The authors created a method that finds these patterns in lots of text, then uses them to decide where questions should be routed, without heavy computation. This new way works just as well as big neural tools but is faster and easier to understand.

What this means in practice

  • For software engineers: Build faster routing systems for language models using interpretable text patterns instead of slow neural embeddings.
  • For customer support teams: Improve automated question routing by applying lightweight, interpretable text feature extraction for speed and reliability.

Authors

Yifan Lu, Qiyue Zhang, Haotian Shan, Hanjie Chen, Jiarong Xing

Abstract

Large Language Model (LLM) routers commonly rely on neural query embeddings, with larger encoders expected to better capture query intent and difficulty. Yet scaling Qwen2.5 encoders from 0.5B to 72B parameters brings little improvement in routing accuracy (Figure 1b), suggesting that small encoders may already capture the query properties needed for routing. We therefore investigate which properties matter and whether they can be extracted directly from text without a neural encoder. We introduce REGEXROUTE, a pipeline that uses sparse autoencoders (SAEs) to discover interpretable regular-expression (regex) features. Using unlabeled text, an LLM turns descriptions of grouped SAE latents into regex extractors and refines them to match latent activation patterns. These extractors supply numerical features to a lightweight routing head, eliminating neural encoding at inference (Figure 1a). Across four benchmarks, one fixed set of 128 features achieves 76.43% average routing accuracy, comparable to 76.41% for the strongest neural text encoder baseline, with much smaller latency and strong robustness. These findings establish explicit, interpretable text features as a practical basis for designing and understanding LLM routers.