End-To-End Multi-Speaker Audio-Visual Automatic Speech Recognition

Patent №

US 11,615,781

Granted

2023-03-28

Filed 2020

Owner

GOOGLE LLC

AI components

5

ml · nlp · vision · speech · kr

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

17062538

A singe audio-visual automated speech recognition model for transcribing speech from audio-visual data includes an encoder frontend and a decoder. The encoder includes an attention mechanism configured to receive an audio track of the audio-visual data and a video portion of the audio-visual data. The video portion of the audio-visual data includes a plurality of video face tracks each associated with a face of a respective person. For each video face track of the plurality of video face tracks, the attention mechanism is configured to determine a confidence score indicating a likelihood that the face of the respective person associated with the video face tack includes a speaking face of the audio track. The decoder is configured to process the audio track and the video face track of the plurality of video face tracks associated with the highest confidence score to determine a speech recognition result of the audio track.

Machine learningNatural languageVisionSpeechKnowledge representationG10L 15/063G10L 15/25G06N 3/0442G06N 3/0455G06N 3/08G06N 3/09G10L 15/16G10L 15/22+1 more

AI classification

Speech1.00
Natural language1.00
Vision1.00
Machine learning1.00
Knowledge representation0.87
AI hardware0.45
Evolutionary computation0.02
Planning0.00

Ownership

GOOGLE LLC

assignment · 539880364

Assignors

BRAGA, OTAVIO

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC