Summary
This paper explores the Soft Mixture of Experts (Soft MoE) architecture, concentrating on its inherent biases and constraints in representational power. Through rigorous theoretical and empirical analyses, the authors show that a single expert within the Soft MoE framework is insufficient to represent even basic convex functions, highlighting the need for multiple experts to achieve optimal performance. The paper presents Theorem 1, which formally substantiates these limitations, and supports it with synthetic data experiments, illustrating that increasing the number of experts enhances performance. Additionally, the authors introduce an efficient method (Algorithm 1: Best Expert Subset Selection) for selecting specialized expert subsets, designed to reduce computational demands in large-scale applications while maintaining model efficacy.
Strengths
- The originality lies in the paper’s creative approach, combining established concepts in expert modeling with a critical examination of representational limitations, as demonstrated by Theorem 1 and the empirical evidence in Appendix B.
- This work has practical significance for advancing the design of scalable MoE models. By addressing the limitations in single-expert representational capacity, the paper opens up discussions on how architectural adjustments, such as multiple experts or efficient selection of subsets, can improve performance and reduce computational costs. These insights are valuable for both researchers and practitioners interested in efficient, scalable solutions for high-dimensional applications in MoE architectures.
Weaknesses
- **Clarity in Methodological Exposition:** Certain aspects of the methodology, such as the approach for selecting the number of experts $k$ for prediction in real applications, would benefit from more detailed and clearer explanations.
- **Clarity in Technical Proof Exposition:** Including a proof sketch at the beginning of Appendix A would enhance readability and provide a clearer roadmap for understanding the technical details. Additionally, a more detailed explanation of why the equalities in lines 711-716 hold would further improve clarity and support the reader's comprehension of the argument. Clarification is needed regarding the notation for norms: are $|\cdot|$ and $\| \cdot \|_2$ intended to represent the same norm?
- **Limited Theoretical Analysis for Multi-Expert Configurations:** Although the paper offers theoretical insights into the limitations of a single expert, it does not extend the theoretical framework to cover multiple-expert $n$ or multiple-slot configurations. A more detailed analysis of how increasing the number of experts impacts scalability, convergence, or approximation accuracy would enhance the strength of the claims presented.
- **Notation Section**: The manuscript would benefit from a dedicated notation section to clarify certain symbols that may be unclear, such as the norm $|\cdot|$ or $\| \cdot \|_2$ for matrices and vectors. Providing a notation section would enhance readability and ensure that readers can follow the mathematical expressions consistently throughout the paper.
- **More Accurate References:** **Lines 040 and 549-550**: The correct citation should be: "Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma (2024). Explaining neural scaling laws. Proceedings of the National Academy of Sciences, 121(27), e2311878121." **Lines 073 and 633-635**: The correct citation should be: "Johan Obando-Ceron, Ghada Sokar, Timon Willi, Clare Lyle, Jesse Farebrother, Jakob Foerster, Gintare Karolina Dziugaite, Doina Precup, and Pablo Samuel Castro. Mixtures of experts unlock parameter scaling for deep RL, 2024. Proceedings of the 41st International Conference on Machine Learning, PMLR 235:38520-38540, 2024." **Lines 073 and 571-574**: The correct citation should be: "Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y.Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models, 2024. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1280–1297, Bangkok, Thailand. Association for Computational Linguistics."
Questions
- **Extension to Multiple Slots per Expert**: From a technical standpoint, are there any challenges preventing the authors from considering only a single slot per expert? If multiple slots $p$, where $p \ge 2$, are considered, would Theorem 1 and the related results still hold? The same question applies to the benign target $t(X) = ||X||_2$: does the result hold only for this specific function, or can it be extended to more general target functions?
- **Multiple Experts Results**: I believe the experiment in Appendix B can be explained by leveraging the approximation properties of functions through mixture of experts models as they relate to the number of experts, as discussed in [1, 2, 3, 4, 5]. At a minimum, the authors should provide a more detailed discussion and/or comparisons with the mentioned references to offer greater insights into how using multiple experts significantly reduces the loss, even when the number of expert parameters in each model is kept constant.
References:
[1] Mendes, E. F., \& Jiang, W. (2012). On convergence rates of mixtures of polynomial experts. Neural computation, 24(11), 3025-3051.
[2] Jiang, W., \& Tanner, M. A. (1999, January). Hierarchical mixtures-of-experts for generalized linear models: some results on denseness and consistency. In Seventh International Workshop on Artificial Intelligence and Statistics. PMLR.
[3] Nguyen, H. D., Lloyd-Jones, L. R., \& McLachlan, G. J. (2016). A universal approximation theorem for mixture-of-experts models. Neural computation, 28(12), 2585-2593.
[4] Nguyen, H. D., Nguyen, T., Chamroukhi, F.,\& McLachlan, G. J. (2021). Approximations of conditional probability density functions in Lebesgue spaces via mixture of experts models. Journal of Statistical Distributions and Applications, 8, 1-15.
[5] Nguyen, H. D., Chamroukhi, F., \& Forbes, F. (2019). Approximation results regarding the multiple-output Gaussian gated mixture of linear experts model. Neurocomputing, 366, 208-214.