Response to Reviewer Y9UT
Dear Reviewer Y9UT,
Thank you for your response and reviewing our work.
---
>Q1: Baselining against SOTA models (SMEAR [1], BTX [2], DeepSeekMoE [3], DEMix [4], etc).
We understand that comparing our work against SOTA MoE models to better understand where SOTA currently stands could indeed enrich our paper. However, the experiments conducted on MH-MoE across models ranging from 300M to 7B parameters were aimed at demonstrating that our proposed MH-MoE can improve the performance of MoE models at various parameter scales, especially when there are many experts, as shown in Figure 7. **This does not imply that our experimental results can be directly compared with the current SOTA models of similar scale (e.g., 7B)**, as these SOTA MoE models differ significantly from our model in terms of training data, training steps, and **especially the scale of activated experts** (e.g., some SOTA models activate 2B parameters, whereas our MH-MoE activates only a few hundred million). Therefore, introducing these results for comparison may not provide meaningful insights.
Furthermore, we want to reiterate that our experiments were conducted under two different MoE frameworks, **with a total of four pretraining tasks and 29 downstream tasks** to validate the effectiveness of MH-MoE. This experimental setup is more extensive than some previous works [1-3], and we believe that our results sufficiently demonstrate the effectiveness of MH-MoE.
**Reference**
[1] Chi, Zewen, et al. "On the representation collapse of sparse mixture of experts." NeurIPS 2022.
[2] Chen, Tianlong, et al. "Adamv-moe: Adaptive multi-task vision mixture-of-experts." ICCV 2023.
[3] Soft Merging of Experts with Adaptive Routing
---
>Q2: Compare with architectures that aimed to solve the same issue.
Thank you for your suggestion. As far as we know, we are the first to identify the issue of low expert activation ratio in MoE models, so we are currently unable to introduce comparisons with similar models in the paper. However, we look forward to more research focusing on this issue in the future, as it is a critical challenge, especially when aiming to increase the number of experts in MoE models significantly. Low expert activation could become a bottleneck limiting the further improvement of MoE models, and we would be very interested in comparing our work with future studies addressing this issue.
---
>Q3: Initial results shared are promising, and I encourage authors to continue investigations.
Thank you for your positive feedback. We will continue our experiments and update the latest results in the upcoming version of the paper.
---
**If you have any further questions or concerns, please feel free to contact us at any time. We are always available and look forward to further discussions with you. :)**
Best regards,
All Authors