Video diffusion model improves with split-role mixture of experts
Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
Video generation models often struggle because they treat all parts of a video the same way, which doesn’t fit video data well. The authors show that existing methods force the model to use all parts evenly, which can break the natural grouping of video information. They propose a new model called SplitMoE that divides experts into two types: one focusing on big-picture meanings and the other on detailed visuals. This helps the model learn faster and create better videos by keeping related information together.
What this means in practice
- •For video generation engineers: Improve video generation quality and training speed by adopting split-role expert models that better organize semantic and visual information.
- •For visual effects teams: Use advanced diffusion techniques with specialized experts to produce more coherent and semantically consistent video content.
Authors
Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang, Yizhi Wang, Xinwei Huang, Minxuan Lin, Angtian Wang, Chongyang Ma, Fan Tang
Abstract
Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and pull-push regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.