Notabletraining methods

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang, Yizhi Wang, Xinwei Huang, Minxuan Lin, Angtian Wang, Chongyang Ma, Fan Tang

Published
Sep 29, 2026 — 17:55 UTC

Problem

Existing visual mixture of experts (MoEs) encounter a uniformity trap, which leads to routing fragmentation and structural distortion. This paper addresses this gap by proposing a new architecture that mitigates these issues, enhancing the performance of video diffusion models. The work is presented as a preprint and has not undergone peer review.

Method

The authors introduce SplitMoE, a split-role sparse architecture designed to improve the efficiency and effectiveness of routing in MoEs. The architecture bifurcates the expert pool into two categories: semantic experts, which focus on specific content understanding, and generic experts, which handle broader tasks. The routing mechanism employed is prototype-guided, allowing for more coherent and contextually relevant expert selection. Additionally, the authors implement a pull-push regularization technique to further enhance the model's performance. Specific details regarding the data used for training and the computational resources required are not disclosed.

Results

The available text does not report quantitative results. However, it is stated that SplitMoE demonstrates superior performance in terms of convergence speed, routing coherence, and video generation quality when compared to traditional load-balanced MoEs.

Limitations

The authors do not report any limitations in their work. However, the lack of detailed data specifications and training compute information may hinder reproducibility and practical application assessments.

Why it matters

The implications of this work are significant for downstream applications in video generation and other areas utilizing MoE architectures. By addressing the uniformity trap, SplitMoE could lead to more efficient models that maintain high-quality outputs, potentially influencing future research directions in sparse architectures and their applications in generative tasks.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI