Looped mixture of experts improve transformer scaling efficiency

Scaling Laws for Looped Mixture of Experts

Machine LearningArtificial IntelligenceComputation and Language

Summary

Transformers are a type of AI model that learn from lots of data but can be expensive to run. This paper studies two ways to make them more efficient: looping (repeating computations) and mixture of experts (using different parts of the model for different tasks). The authors created mathematical scaling laws that combine these two ideas, allowing better predictions of model performance with limited compute and memory. Their results show that combining looping and sparsity leads to improved reasoning ability and more efficient use of resources compared to using either method alone.

What this means in practice

  • For machine learning engineers: Design transformer models that use looping and mixture of experts to optimize performance within fixed compute and memory limits.
  • For ai research infrastructure teams: Deploy scalable training pipelines that incorporate looped mixture of experts to improve parameter efficiency and handle large-scale reasoning tasks.

Authors

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

Abstract

Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.