METHOD AND SYSTEM FOR ENHANCING A SPEECH SIGNAL OF A HUMAN SPEAKER IN A VIDEO USING VISUAL INFORMATION

Patent №

US 10,475,465

Granted

2019-11-12

Filed 2018

Owner

YISSUM RESEARCH DEVELOPMENT COMPANY OF THE HEBREW UNIVERSITY OF JERUSALEM LTD.

Lab

AI components

4

ml · nlp · vision · speech

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16026449

A method and system for enhancing a speech signal is provided herein. The method may include the following steps: obtaining an original video, wherein the original video includes a sequence of original input images showing a face of at least one human speaker, and an original soundtrack synchronized with said sequence of images; and processing, using a computer processor, the original video, to yield an enhanced speech signal of said at least one human speaker, by detecting sounds that are acoustically unrelated to the speech of the at least one human speaker, based on visual data derived from the sequence of original input images.

Machine learningNatural languageVisionSpeechG06V 10/82G06F 18/2413G06F 18/251G06V 10/454G06V 10/764G06V 10/803G06V 40/161G10L 21/0216+8 more

AI classification

Speech1.00
Vision1.00
Machine learning0.94
Natural language0.93
AI hardware0.02
Planning0.00
Evolutionary computation0.00
Knowledge representation0.00

Ownership

YISSUM RESEARCH DEVELOPMENT COMPANY OF THE HEBREW UNIVERSITY OF JERUSALEM LTD.

assignment · 470870262

Assignors

PELEG, SHMUEL, SHAMIR, ASAPH, HALPERIN, TAVI, GABBAY, AVIV, EPHRAT, ARIEL

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC