DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning

The Mixture-of-Experts (MoE) architecture is showing promising results in\nimproving parameter sharing in multi-task learning (MTL) and in scaling\nhigh-capacity neural networks. State-of-the-art MoE models use a trainable\nsparse gate to select a subset of the experts for each input example. While\nconceptually appealing, existing sparse gates, such as Top-k, are not smooth.\nThe lack of smoothness can lead to convergence and statistical performance\nissues when training with gradient-based methods. In this paper, we develop\nDSelect-k: a continuously differentiable and sparse gate for MoE, based on a\nnovel binary encoding formulation. The gate can be trained using first-order\nmethods, such as stochastic gradient descent, and offers explicit control over\nthe number of experts to select. We demonstrate the effectiveness of DSelect-k\non both synthetic and real MTL datasets with up to $128$ tasks. Our experiments\nindicate that DSelect-k can achieve statistically significant improvements in\nprediction and expert selection over popular MoE gates. Notably, on a\nreal-world, large-scale recommender system, DSelect-k achieves over $22\\%$\nimprovement in predictive performance compared to Top-k. We provide an\nopen-source implementation of DSelect-k.\n

Paper

References (45)

Scroll for more · 33 remaining

Similar papers

© 2026 NYSGPT2525 LLC