How to Loop MoE: Flatten the Experts, Untie the Attention
Shouren Wang, Chuang Ma, Mohsen Hariri, Debargha Ganguly, Wang Yang, Xiaoqing Tong, Qianying Liu, Xiaotian Han, Vipin Chaudhary
- Published
- Sep 28, 2026 — 17:58 UTC
Problem
The paper addresses a gap in the effective utilization of experts in mixture of experts (MoE) models, particularly focusing on the limitations of existing architectures in optimizing expert parameters and attention mechanisms. The authors propose a new approach to enhance the performance of MoE models, which is particularly relevant given the increasing complexity and scale of modern neural networks. This work is presented as a preprint and has not yet undergone peer review.
Method
The authors introduce the Looped MoE architecture with a novel component called Foil. Key features of this architecture include:
- Expert Parameters: The parameters of the experts are fixed, which simplifies the training process and reduces computational overhead.
- Expert Compute per Token: The compute required per token is also fixed, ensuring consistent resource allocation during training.
- Flattening Experts: The architecture halves the number of expert layers while doubling the number of experts per layer and increasing the number of passes through the model. This design aims to enhance the efficiency of expert utilization.
- Untying Attention: Each pass through the model utilizes unique attention parameters, allowing for more flexible and adaptive attention mechanisms. The experts and routers are shared across passes, which contributes to improved routing efficiency.
Results
The results demonstrate the effectiveness of the proposed Looped MoE architecture:
- Pretraining Loss: The Foil model achieves a lower pretraining loss compared to the baseline at 20 billion tokens.
- Pretraining Loss Improvement: At 100 billion tokens, the model shows an improvement of 0.012 nat below the baseline loss, indicating enhanced training efficiency with maximum flattening.
- Downstream Accuracy: The accuracy on downstream tasks is reported to be on par with or better than the baseline models, suggesting that the architectural changes do not compromise performance.
- Routing Confidence: The routing mechanism exhibits more balanced and confident routing, maintaining an equal shape across the model, which is crucial for effective expert utilization.
Limitations
The authors do not report any limitations in their work. However, the absence of reported limitations may suggest a need for further empirical validation in diverse settings or with different datasets to fully assess the robustness of the proposed architecture.
Why it matters
This work has significant implications for the design of future MoE models, particularly in large-scale applications where efficient expert utilization and adaptive attention mechanisms are critical. By optimizing the structure and functionality of MoE architectures, this research could lead to advancements in model performance and resource efficiency, paving the way for more scalable and effective AI systems.
By Callan Zhang · Sep 28, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
