EFFICIENT TRANSFORMER LANGUAGE MODELS WITH DISENTANGLED ATTENTION AND MULTI-STEP DECODING

Patent №

US 11,526,679

Granted

2022-12-13

Filed 2020

Owner

MICROSOFT TECHNOLOGY LICENSING, LLC

AI components

5

ml · nlp · vision · speech · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16910508

Systems and methods are provided for facilitating the building and use of natural language understanding models. The systems and methods identify a plurality of tokens and use them to generate one or more pre-trained natural language models using a transformer. The transformer disentangles the content embedding and positional embedding in the computation of its attention matrix. Systems and methods are also provided to facilitate self-training of the pre-trained natural language model by utilizing multi-step decoding to better reconstruct masked tokens and improve pre-training convergence.

Machine learningNatural languageVisionSpeechAI hardwareG06F 40/40G06N 3/088G06F 40/237G06F 40/30G06N 3/045G06N 3/0455G06N 3/048G06N 3/0499+2 more

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
Vision1.00
AI hardware1.00
Planning0.02
Evolutionary computation0.01
Knowledge representation0.00

Ownership

MICROSOFT TECHNOLOGY LICENSING, LLC

assignment · 530280435

Assignors

HE, PENGCHENG, LIU, XIAODONG, GAO, JIANFENG, CHEN, WEIZHU

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC