GENERATING RESPONSES TO QUERIES ABOUT VIDEOS UTILIZING A MULTI-MODAL NEURAL NETWORK WITH ATTENTION

Patent №

US 11,615,308

Granted

2023-03-28

Filed 2021

Owner

ADOBE INC.

Lab

AI components

7

ml · nlp · vision · speech · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

17563901

The present disclosure relates to systems, methods, and non-transitory computer-readable media for generating a response to a question received from a user during display or playback of a video segment by utilizing a query-response-neural network. The disclosed systems can extract a query vector from a question corresponding to the video segment using the query-response-neural network. The disclosed systems further generate context vectors representing both visual cues and transcript cues corresponding to the video segment using context encoders or other layers from the query-response-neural network. By utilizing additional layers from the query-response-neural network, the disclosed systems generate (i) a query-context vector based on the query vector and the context vectors, and (ii) candidate-response vectors representing candidate responses to the question from a domain-knowledge base or other source. To respond to a user's question, the disclosed systems further select a response from the candidate responses based on a comparison of the query-context vector and the candidate-response vectors.

Machine learningNatural languageVisionSpeechKnowledge representationPlanningAI hardwareG06N 3/08G06F 17/16G06N 3/02G06N 3/044G06N 3/0442G06N 3/045G06N 3/0455G06N 3/0464+8 more

AI classification

Natural language1.00
Machine learning1.00
Vision1.00
Speech1.00
AI hardware1.00
Planning0.99
Knowledge representation0.98
Evolutionary computation0.00

Ownership

ADOBE INC.

assignment · 584930453

Assignors

ZHAO, WENTIAN, KIM, SEOKHWAN, XU, NING, JIN, HAILIN

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC