Papers for

ai research infrastructure teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Looped mixture of experts improve transformer scaling efficiency

Scaling Laws for Looped Mixture of Experts

Abstract: Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mapping that characterizes the effective-parameter gain from looping and how sparsity raises this gain. The laws predict the held-out loss of looped models more accurately than prior alternatives, and recover the standard dense and MoE scaling laws as special cases. Beyond prediction, the fitted laws provide a principled foundation for designing looped MoE models under compute and memory constraints. Downstream evaluations further demonstrate the complementary benefits of the two axes: sparsity delivers ~3x active-parameter efficiency, recurrence yields ~2x total-parameter efficiency on reasoning, and joint scaling further advances the performance frontier. As a practical extension, we show these gains hold at trillion-token scale: at matched training compute, a looped MoE with law-derived recurrence matches a ~2x larger non-looped MoE on the reasoning benchmarks, while enabling test-time scaling through recurrence.

Wed 30 SeptMachine LearningArtificial IntelligenceComputation and Language
The gist
Transformers are a type of AI model that learn from lots of data but can be expensive to run. This paper studies two ways to make them more efficient: looping (repeating computations) and mixture of experts (using different parts of the model for different tasks). The authors created mathematical scaling laws that combine these two ideas, allowing better predictions of model performance with limited compute and memory. Their results show that combining looping and sparsity leads to improved reasoning ability and more efficient use of resources compared to using either method alone.
Open → 2609.40316v1

Stabilizing reinforcement learning in large language models for better reasoning

Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification

Abstract: While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from "Rollout Silencing" and low-quality gradient signals in standard sampling procedures. In this work, we propose a robust, data-centric framework to stabilize RL training. We first introduce Potential-Aware Query Mining (PAQM), which filters data dynamically to focus on the "Distillation Zone"---samples with high potential for capability elicitation. Furthermore, we present Hybrid Stratified Replay (HSR), a novel mechanism that restructures batches by stratifying rollouts based on Path Entropy, a rollout-level confidence proxy, and outcome reward. Within each optimization step, HSR reuses current-policy "Stability Anchors" and "Hard Negatives" to construct high-contrast optimization groups, then clears its buffers before the next step. This approach mitigates entropy collapse while improving the utilization of learning signals under limited compute. Our method outperforms strong baselines on complex reasoning tasks, offering a principled solution for stable and efficient RL fine-tuning.

Mon 7 SeptMachine LearningComputer Vision and Pattern Recognition
The gist
Training large language models to think step-by-step can be unstable and confusing for the system, causing poor learning signals. The authors introduce methods that carefully select the right training examples and organize learning batches by confidence and reward to maintain balance. This approach helps the model learn more steadily and effectively, especially on tricky reasoning tasks. These techniques improve the training process without needing extra compute power.
Open → 2609.07148v1