Patent №
US 11,288,438
Granted
2022-03-29
Filed 2020
Owner
SALESFORCE.COM, INC.
Lab
—
AI components
6
ml · nlp · vision · speech · kr · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
16781223
Systems and methods are provided for performing a video-grounded dialogue task by a neural network model using bi-directional spatial-temporal reasoning. According to some embodiments, the systems and methods implement a dual network architecture or framework. This framework includes one network or reasoning module that learns dependencies between text and video in the direction of spatial→temporal, and another network or reasoning module that learns in the direction of temporal→spatial. The output of the multimodal reasoning modules may be combined to learn dependencies between language features in dialogues. The result joint representation is used as a contextual feature to the decoding components which allow the model to semantically generate meaningful responses to the users. In some embodiments, pointer networks are extended to the video-grounded dialogue task to allow the model to point to specific tokens from multiple source sequences to generate responses.
AI classification
Ownership
SALESFORCE.COM, INC.
assignment · 518580756
Assignors
LE, HUNG, HOI, CHU HONG
On an employer assignment, the assignors are typically the inventors.