SPEECH DETECTION AND ENHANCEMENT USING AUDIO/VIDEO FUSION

Patent №

US 7,689,413

Granted

2010-03-30

Filed 2007

Owner

MICROSOFT CORPORATION

AI components

5

ml · nlp · vision · speech · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11852961

A system and method facilitating speech detection and/or enhancement utilizing audio/video fusion is provided. The present invention fuses audio and video in a probabilistic generative model that implements cross-model, self-supervised learning, enabling rapid adaptation to audio visual data. The system can learn to detect and enhance speech in noise given only a short (e.g., 30 second) sequence of audio-visual data. In addition, it automatically learns to track the lips as they move around in the video.

AI classification

Vision1.00
Speech1.00
Machine learning1.00
Natural language1.00
AI hardware0.64
Knowledge representation0.03
Evolutionary computation0.00
Planning0.00

Ownership

MICROSOFT CORPORATION

assignment · 198750498

Assignors

HERSHEY, JOHN R., KRISTJANSSON, TRAUSTI THOR, ATTIAS, HAGAI, JOJIC, NEBOJSA

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC