SYSTEMS AND METHODS FOR USING LATENT VARIABLE MODELING FOR MULTI-MODAL VIDEO INDEXING
Patent №
US 9,542,934
Granted
2017-01-10
Filed 2014
Owner
FUJI XEROX CO., LTD.
Lab
—
AI components
5
nlp · vision · speech · kr · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
14192861
A computer-implemented method performed in connection with a computerized system incorporating a processing unit and a memory, the computer-implemented method involving: using the processing unit to generate a multi-modal language model for co-occurrence of spoken words and displayed text in the plurality of videos; selecting at least a portion of a first video; extracting a plurality of spoken words from the selected portion of the first video; extracting a first displayed text from the selected portion of the first video; and using the processing unit and the generated multi-modal language model to rank the extracted plurality of spoken words based on probability of occurrence conditioned on the extracted first displayed text.
AI classification
Ownership
FUJI XEROX CO., LTD.
assignment · 323860709
Assignors
COOPER, MATTHEW L., JOSHI, DHIRAJ, CHEN, HUIZHONG
On an employer assignment, the assignors are typically the inventors.