MIXTURE OF EXPERTS MODELS WITH SPARSIFIED WEIGHTS

Patent №

US 12,694,265

Granted

2026-07-28

Filed 2022

Owner

Microsoft Technology Licensing, LLC

AI components

0

Assignment

None on record

Dataset

AIPD

Application

17657604

A method is presented for operating a machine learning model including one or more mixture of experts layers. The method comprises receiving one or more input data shards at a routing gate network for a mixture of experts layer comprising a plurality of neural network experts. One or more neural network experts in the mixture of experts layer is designated layer to evaluate each input data shard. For each designated neural network expert, a weight matrix is retrieved having a predetermined sparsity to generate a sparsified designated neural network expert. Each input data shard is evaluated with a respective sparsified designated neural network expert.

G06N 3/042G06N 3/082G06N 3/0495G06N 3/045G06N 3/092G06N 3/08G06N 3/084

Ownership

Microsoft Technology Licensing, LLC

© 2026 NYSGPT2525 LLC