BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models
Gang Fu, Adel Javanmard, MohammadHossein Bateni, Vahab Mirrokni
- Published
- Oct 5, 2026 — 17:15 UTC
Problem
This work addresses the issue of imbalanced expert utilization in conventional flat routers for Mixture-of-Experts (MoE) layers. The authors highlight that existing routing mechanisms often lead to suboptimal performance due to uneven load distribution among experts, which can hinder the efficiency and effectiveness of large embedding models. This paper is a preprint and has not undergone peer review.
Method
The proposed architecture, BRANCH-MoE, employs a binary decision tree structure with a total of E experts located at the leaves. The routing mechanism utilizes a branching probability that is determined by the arrival-weighted mean score at the internal nodes of the tree. To estimate this mean score, the authors implement an exponential moving average, which allows for a dynamic adjustment based on incoming data. Notably, BRANCH-MoE does not require an auxiliary load-balancing loss, simplifying the training process. The authors also discuss a noise-lag trade-off that arises from the moving-average estimate, which can impact the routing decisions. Furthermore, the frequency of expert execution is shown to influence the convergence rate of stochastic gradient descent, while confident decisions made near the root of the tree help to minimize cross-device communication overhead.
Results
The available text does not report quantitative results. However, it asserts that hierarchical routing in BRANCH-MoE preserves task quality while achieving balanced utilization when compared to several baselines, including Switch softmax, DeepSeek-V3 dynamic-bias, Skywork logit-normalized, and deterministic hash routing.
Limitations
The authors do not report any limitations in their work. However, the absence of quantitative results may limit the ability to fully assess the performance improvements over existing methods.
Why it matters
The implications of this work are significant for the development of more efficient MoE architectures, particularly in scenarios where expert utilization is critical for performance. By addressing the imbalance in expert routing, BRANCH-MoE could lead to advancements in large-scale models, enhancing their scalability and effectiveness in various applications.
By Turing Wire Research Desk · Oct 5, 2026 · How we work →
Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.
Source: arXiv cs.AI
