Convergence Rates for Softmax Gating Mixture of Experts

Mixture of experts (MoE) has recently emerged as an effective framework for deploying machine learning models in a scalable and efficient way by softly dividing complex tasks among multiple specialized sub-models termed experts. Central to the success of MoE is an adaptive gating mechanism which determines the relevance of each expert to a given input and then dynamically assigns experts their respective weights. Despite its widespread use in practice, a comprehensive study on the effects of the softmax gating on the MoE has been lacking in the literature. To bridge this gap, we conduct a thorough theoretical analysis of the convergence rates for the problem of parameter estimation and expert estimation. We consider standard softmax gating and several variants, including a dense-to-sparse gating and a hierarchical softmax gating. Our theoretical results provide useful insights into the design of sample-efficient expert structures. In particular, we demonstrate that it requires polynomially many data points to estimate experts satisfying our proposed strong identifiability condition, namely a commonly used two-layer feed-forward network. In stark contrast, estimating linear experts, which violate the strong identifiability condition, necessitates exponentially many data points as a result of intrinsic parameter interactions, which we express in the language of partial differential equations.

Paper

Similar papers

© 2026 NYSGPT2525 LLC