Language models trained to run well at any compute budget
Telescopic Language Models
Computation and LanguageArtificial Intelligence
Summary
Language models usually need different training or compression steps for different amounts of computing power. The authors trained a single model, called a telescopic language model, that works well whether it runs a little or a lot. They do this by training the model so that any smaller version of it still predicts language reliably without needing changes at use time. This approach reduces training costs and provides a smooth range of compute-versus-performance options. Their results show the key is the training method, not just the model’s design.
What this means in practice
- •For machine learning engineers: Train and deploy a single language model that adapts smoothly to different compute constraints without retraining or separate compression.
- •For cloud service providers: Offer flexible AI language model services that adjust computation dynamically for different user budgets using one trained model artifact.$Commercial implications: Enables selling scalable language models that optimize costs and performance per customer without separate model instances.
Authors
Zhilin Guo, Boqiao Zhang, Hakan Aktas, Kyle Fogarty, Nursena Koprucu Aslan, Wenzhao Li, Canberk Baykal, Albert Miao, Siyu Hong, Yixiao Liu, Adam Wu, Ashish Kumar Singh, Sakar Khattar, Chenliang Zhou, Weihao Xia, Cristina Nader Vasconcelos, Cengiz Oztireli
Abstract
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested model at chance level everywhere else (perplexity 10^2-10^5 in our baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data stream for all methods), a single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks, reducing the area under the quality-budget curve by 43-44% relative to the fixed-exit suites while matching them at full capacity, at ~12% lower GPU cost per run. The prefix sampling density is a dial: concentrating it on a few depths recovers fixed-exit quality there at the price of the continuum, so the operating points become a training-time choice rather than an architectural one. These results indicate that the training objective, not the nesting itself, is what makes a model elastic.