Bi-Directional Spatial-Temporal Reasoning for Video-Grounded Dialogues

Patent №

US 11,288,438

Granted

2022-03-29

Filed 2020

Owner

SALESFORCE.COM, INC.

Lab

AI components

6

ml · nlp · vision · speech · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16781223

Systems and methods are provided for performing a video-grounded dialogue task by a neural network model using bi-directional spatial-temporal reasoning. According to some embodiments, the systems and methods implement a dual network architecture or framework. This framework includes one network or reasoning module that learns dependencies between text and video in the direction of spatial→temporal, and another network or reasoning module that learns in the direction of temporal→spatial. The output of the multimodal reasoning modules may be combined to learn dependencies between language features in dialogues. The result joint representation is used as a contextual feature to the decoding components which allow the model to semantically generate meaningful responses to the users. In some embodiments, pointer networks are extended to the video-grounded dialogue task to allow the model to point to specific tokens from multiple source sequences to generate responses.

Machine learningNatural languageVisionSpeechKnowledge representationAI hardwareG06F 40/10G06F 40/30G06F 40/284G06N 3/044G06N 3/045G06N 3/0455G06N 3/0464G06N 3/0475+5 more

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
AI hardware1.00
Vision1.00
Knowledge representation0.99
Evolutionary computation0.00
Planning0.00

Ownership

SALESFORCE.COM, INC.

assignment · 518580756

Assignors

LE, HUNG, HOI, CHU HONG

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC