Efficient Memory Transformer Based Acoustic Model for Low Latency Streaming Speech Recognition

Patent №

US 11,646,017

Granted

2023-05-09

Filed 2021

Owner

FACEBOOK, INC.

AI components

7

ml · nlp · vision · speech · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

17193414

In one embodiment, a method includes accessing a machine-learning model configured to generate an encoding for an utterance by using a module to process data associated with each segment of the utterance in a series of iterations, performing operations associated with an i-th segment during an n-th iteration by the module, which include receiving an input comprising input contextual embeddings generated for the i-th segment in a preceding iteration and a memory bank storing memory vectors generated in the preceding iteration for segments preceding the i-th segment, generating attention outputs and a memory vector based on keys, values, and queries generated using the input, and generating output contextual embeddings for the i-th segment based on the attention outputs, providing the memory vector to the module for performing operations associated with the i-th segment in a next iteration, and performing speech recognition by decoding the encoding of the utterance.

Machine learningNatural languageVisionSpeechKnowledge representationPlanningAI hardwareG10L 15/183G06N 3/044G06N 3/045G06N 3/08G06N 3/084G10L 15/16G10L 15/22G10L 15/28

AI classification

Speech1.00
Natural language1.00
Machine learning1.00
AI hardware1.00
Planning1.00
Knowledge representation0.96
Vision0.78
Evolutionary computation0.01

Ownership

FACEBOOK, INC.

assignment · 561330978

Assignors

SHI, YANGYANG, WANG, YONGQIANG, WU, CHUNYANG, YEH, CHING-FENG, CHAN, JULIAN YUI-HIN, ZHANG, QIAOCHU, LE, DUC HOANG, SELTZER, MICHAEL LEWIS

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC