Papers for

cloud ai infrastructure teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Looped mixture of experts improves model efficiency by reshaping experts

How to Loop MoE: Flatten the Experts, Untie the Attention

Abstract: Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per token held fixed, Foil (1) flattens the experts, halving the expert layers, doubling the experts per layer and doubling the passes, so that every routing decision chooses from a larger pool, and (2) unties the attention, giving each pass its own attention parameters while the experts and routers stay shared. Experiments show that Foil clearly outperforms the unflattened looped baseline: at 20B tokens every Foil model has lower pretraining loss than the baseline; at 100B tokens the loss improves monotonically with the degree of flattening, the most flattened Foil ending 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better; untying the attention also yields more balanced and more confident routing at equal shape. Our ablations analyse why Foil works and turn the findings into design guidance for looped MoE: the returns of looping and of widening the expert layers amplify each other, routing confidence tracks healthy expert use better than load balance, and a sparse looped MoE should therefore use more experts per layer and more passes. Code and configurations are available at https://github.com/SR-A-W/how-to-loop-moe.

Mon 28 SeptMachine LearningArtificial IntelligenceComputation and Language
The gist
Large AI models sometimes reuse parts of themselves repeatedly to do more with less. This paper shows a way to reshape and reorganize these reused parts, called experts, so the system can choose from a bigger and better pool each time. The authors found that this approach leads to better learning and more balanced use of the parts, without increasing the model size or computation needed. This means AI models can perform better by simply looping through a clever arrangement of their own components.
Open → 2609.35751v1

Kv cache compression method retains long and off context query strength

Cartridges++: KV Cache Compression without Off-Context Derailment

Abstract: Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all, ahead of inference time. Methods to obtain CKVs range from drop mechanisms that reduce their number of columns, to learned approaches. Among the latter, Cartridges have emerged as a leading compression method, learning compact KV representations through distillation on relevant Q/A pairs. While existing evaluations focus primarily on whether Cartridges and other CKVs yield approximately similar responses to document-related, on-context queries, we investigate the crucial deployment question of whether they can handle off-context queries, something the native KV representation is particularly good at, thanks to the mechanics of attention. We observe a fundamental trade-off: while Cartridges perform better for on-context queries, heuristic-variants preserve better the original LLM's ability to operate off-context. We measure this through their capability to avoid context contamination in their response, retain general knowledge, and follow instructions. We propose Cartridges++, simple modifications to cartridges that retain off-context abilities at small or negligible cost. The router variant decides at inference time whether the query should use the learned long-context memory, while the data-mixing variant allocates a small fraction of training Q/As to queries outside the reference long document. Our study shows that assessing CKVs on document utility alone can mask substantial degradation in broader model capabilities, yet those issues can be fixed with benign changes to CKV inference or training.

Mon 28 SeptMachine Learning
The gist
Working with long documents on large language models uses a lot of computer memory and time, so people try to compress the memory it needs. The authors study a popular compression method called Cartridges, which works well for questions about the original document but struggles when questions are unrelated to that document. They find a trade-off between remembering document details and handling general queries. They propose Cartridges++, which modify the compression to keep the benefits for document queries while also preserving the model’s ability to answer off-topic questions well.
Open → 2609.35621v1

Efficiently select experts to speed up large AI models

EAT: Expert Account Tracker for Efficient MoE Inference

Abstract: Mixture-of-Experts (MoE) models have emerged as a revolutionary method to scale Transformer models. However, traditional MoE architecture still suffers from inefficiency since a large number of experts are unnecessarily activated. Existing approaches for reducing the number of activated experts often overlook the historical performance of each expert. In this paper, we propose EAT, a novel method called Expert Account Tracker (EAT), which utilizes history-awareness metrics and adaptive thresholding to dynamically select the most important experts, thereby reducing the activated expert number while effectively maintaining the model performance. Experiments show that EAT outperforms the existing baseline Top-P method across multiple models and datasets, achieving over 25% an average reduction compared to the vanilla method in the number of activated experts and performing better token generation speed compared to the baseline. Furthermore, the performance of pruned models can be efficiently recovered via OPD using only 9K data. Additionally, through ablation studies, we find that excessively reducing the number of activated experts can significantly harm model performance, and the importance of experts varies across layers, with higher-level experts being generally more critical.

Sun 27 SeptArtificial Intelligence
The gist
Large AI models called Mixture-of-Experts (MoE) use many smaller modules called experts to make decisions, but activating too many experts wastes time and resources. The authors introduced EAT, a method that keeps track of how well each expert has performed in the past to smartly choose only the most useful ones for each task. This reduces the number of experts that need to be activated by over 25%, making the model faster without losing accuracy. They also show that the model’s performance can be recovered quickly if experts are pruned too much.
Open → 2609.33614v1

Oled moe speeds up large language model inference with smarter expert memory use

OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading

Abstract: Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a large expert parameter footprint that exceeds memory-constrained GPU capacity. Expert offloading is a natural remedy, yet existing MoE serving systems target autoregressive decoding and rely on intra-iteration layer-wise prefetching: while computing one layer, they predict and load experts for subsequent layers. Under dLLM inference, block-wise routing expands the active expert working set within each iteration, making such prefetches difficult to complete in time and costly when mispredicted. Consequently, existing prefetch-based solutions often degenerate into on-demand expert loading with high decoding latency. We propose OLED-MoE, an expert offloading system that shifts the optimization target from intra-iteration prefetching to inter-iteration expert retention. Its key insight is that adjacent denoising iterations exhibit strong expert routing overlap, and token confidence indicates which experts are likely to be reused. OLED-MoE uses confidence-guided inter-iteration prediction to retain high-value experts in GPU memory without introducing extra prefetch traffic. It further compensates unavoidable cache misses through CPU-GPU cooperative execution, jointly considering dynamic expert computation load and predicted future reuse. Across diverse dLLM workloads, OLED-MoE reduces time per output token (TPOT) by 1.23x-7.93x and improves expert cache utilization by 1.44x-4.23x over state-of-the-art offloading systems. Notably, OLED-MoE approaches full-residency performance while using only 40% of the expert GPU memory, incurring merely 23% higher TPOT despite a 60% reduction in expert memory footprint. OLED-MoE's source code is publicly available at https://github.com/flashserve/OLED-MoE.

Sun 27 SeptDistributed, Parallel, and Cluster ComputingComputation and Language
The gist
Large language models that use many expert components can become too big to fit into standard graphics cards, slowing down their response time. The authors found existing methods for managing these experts during decoding often miss timing targets, causing delays. They designed OLED-MoE, a system that smartly keeps frequently reused experts in fast memory across consecutive steps, using token confidence to predict reuse. This approach reduces waiting time significantly and uses memory much more efficiently while keeping performance close to ideal full-memory use.
Open → 2609.33385v1

On policy training improves efficient transformer attention on long tasks

On-Policy Attention Linearization

Abstract: Hybrid transformer architectures that replace most softmax attention layers with linear attention offer transformer-level quality at a fraction of the memory cost. Rather than pretraining such models, a growing body of work distills them from already trained full-attention transformers. However, these distilled models often collapse on long-context retrieval and reasoning tasks, particularly when operating in thinking mode, where the efficiency gains of hybrid architectures matter most. Since linear attention layers must compress context into a fixed-size state, their errors compound over long sequences. As off-policy distillation never teaches the student model to recover from this drift, tasks that necessitate longer sequence lengths become especially challenging. We introduce On-Policy Attention Linearization (OPAL) in which the hybrid attention student samples its own long-context trajectories and receives dense supervision from the frozen full-attention teacher. Applying OPAL to Qwen3-4B and MiMo-7B-RL-0530, we recover $87$--$94\%$ of full-attention performance on commonsense reasoning, $100\%$ on needle-in-a-haystack (NIAH) retrieval, and $83$--$93\%$ on mathematical reasoning with only 3B training tokens. We achieve these results without supervised fine-tuning (SFT) or reinforcement learning with verifiable rewards (RLVR). Compared with the strongest prior linearization method, which recovers $68\%$ of its teacher's retrieval performance and $21.6\%$ absolute average mathematical reasoning accuracy, OPAL fully recovers retrieval and achieves $67.6$--$72.2\%$ on math reasoning.

Fri 25 SeptMachine Learning
The gist
Transformers often use a type of attention called softmax attention that requires a lot of memory and computing power, so some newer models use linear attention to be faster and use less memory. However, these faster versions often struggle with tasks that need understanding of long pieces of information, especially in complex reasoning. The authors introduced a new way to train the faster model by letting it learn from its own long attempts while being guided by a bigger model. This method greatly improves performance on tasks like reasoning, finding tiny details, and math problems without needing extra complex training.
Open → 2609.31947v1

Hybrid sparse attention improves long-context processing efficiency

HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

Abstract: Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate long-context retrieval. To meet these demands, we introduce HySparse2, a hybrid sparse attention architecture with two-level KV sharing. At the outer level, KV Bridging adopts a YOCO-style self-decoder and cross-decoder structure, but bridges only full-attention layers. The self-decoder uses hybrid sliding-window attention (SWA), while the cross-decoder uses hybrid sparse attention. The KV caches for full-attention layers in the cross-decoder are generated from the hidden states of full-attention layers in the self-decoder. At the inner level, HySparse2 retains HySparse's core KV Reuse design with two refinements. First, it replaces block-level sparsity with token-level sparsity for finer long-context retrieval. Second, it removes the separate SWA branch from sparse layers and instead forces a sliding window of recent tokens into the sparse selection. This two-level KV sharing allows all cross-decoder KV caches to be constructed from self-decoder hidden states. Prefill can therefore exit after the self-decoder, skipping all cross-decoder layers. On an 80B-A3B MoE model, HySparse2 outperforms HySparse and Hybrid SWA on long-context retrieval and multi-turn agentic tasks, while substantially reducing prefill computation and KV-cache storage.

Tue 22 SeptComputation and Language
The gist
Long-running AI agents often need to understand and remember lots of information while working on many small tasks. The authors introduced HySparse2, a new way for AI models to handle this big amount of information more efficiently. It combines different attention methods and shares intermediate data cleverly to save memory and speed up processing. Their approach helps AI remember more details over longer conversations or actions while using less computing power.
Open → 2609.26368v1