Mixture-of-Experts under Finite-Rate Gating: Communication--Generalization Trade-offs

Mixture-of-Experts (MoE) architectures decompose prediction tasks into specialized expert sub-networks selected by a gating mechanism. This letter adopts a communication-theoretic view of MoE gating, modeling the gate as a stochastic channel operating under a finite information rate. Within an information-theoretic learning framework, we specialize a mutual-information generalization bound and develop a rate-distortion characterization <inline-formula> <tex-math notation="LaTeX">$D(R_{g})$ </tex-math></inline-formula> of finite-rate gating, where <inline-formula> <tex-math notation="LaTeX">$R_{g}:=I(X; T)$ </tex-math></inline-formula>, yielding (under a standard empirical rate-distortion optimality condition) <inline-formula> <tex-math notation="LaTeX">$\mathbb {E}[R(W)] \le D(R_{g})+\delta _{m}+\sqrt {(2/m){\,}I(S; W)}$ </tex-math></inline-formula>. The analysis yields capacity-aware limits for communication-constrained MoE systems, and numerical simulations on synthetic multi-expert models empirically confirm the predicted trade-offs between gating rate, expressivity, and generalization.

Paper

Similar papers

© 2026 NYSGPT2525 LLC