METHODS AND APPARATUS FOR DECISION FUSION OF AUDIO AND VIDEO BASED SPEAKER IDENTIFICATION FOR MULTIMEDIA INFORMATION ACCESS

Patent №

US 6,567,775

Granted

2003-05-20

Filed 2000

Owner

INTERNATIONAL BUSINESS MACHINES CORPORATION

AI components

4

ml · nlp · vision · speech

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

09558371

A method and apparatus are disclosed for identifying a speaker in an audio-video source using both audio and video information. An audio-based speaker identification system identifies one or more potential speakers for a given segment using an enrolled speaker database. A video-based speaker identification system identifies one or more potential speakers for a given segment using a face detector/recognizer and an enrolled face database. An audio-video decision fusion process evaluates the individuals identified by the audio-based and video-based speaker identification systems and determines the speaker of an utterance in accordance with the present invention. A linear variation is imposed on the ranked-lists produced using the audio and video information. The decision fusion scheme of the present invention is based on a linear combination of the audio and the video ranked-lists. The line with the higher slope is assumed to convey more discriminative information. The normalized slopes of the two lines are used as the weight of the respective results when combining the scores from the audio-based and video-based speaker analysis. In this manner, the weights are derived from the data itself.

AI classification

Vision1.00
Speech1.00
Machine learning1.00
Natural language0.61
Knowledge representation0.13
Planning0.02
Evolutionary computation0.01
AI hardware0.00

Ownership

INTERNATIONAL BUSINESS MACHINES CORPORATION

assignment · 110800354

Assignors

MAALI, FEREYDOUN, VISWANATHAN, MAHESH

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC