SYSTEMS AND METHODS FOR USING LATENT VARIABLE MODELING FOR MULTI-MODAL VIDEO INDEXING

Patent №

US 9,542,934

Granted

2017-01-10

Filed 2014

Owner

FUJI XEROX CO., LTD.

Lab

AI components

5

nlp · vision · speech · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

14192861

A computer-implemented method performed in connection with a computerized system incorporating a processing unit and a memory, the computer-implemented method involving: using the processing unit to generate a multi-modal language model for co-occurrence of spoken words and displayed text in the plurality of videos; selecting at least a portion of a first video; extracting a plurality of spoken words from the selected portion of the first video; extracting a first displayed text from the selected portion of the first video; and using the processing unit and the generated multi-modal language model to rank the extracted plurality of spoken words based on probability of occurrence conditioned on the extracted first displayed text.

Natural languageVisionSpeechKnowledge representationAI hardwareG10L 15/05G06F 16/7834G06V 20/635G11B 27/00G11B 27/28H04N 21/234336H04N 21/440236G06N 7/01+2 more

AI classification

Natural language1.00
Speech1.00
Vision1.00
Knowledge representation0.98
AI hardware0.93
Machine learning0.42
Planning0.03
Evolutionary computation0.01

Ownership

FUJI XEROX CO., LTD.

assignment · 323860709

Assignors

COOPER, MATTHEW L., JOSHI, DHIRAJ, CHEN, HUIZHONG

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC