Spatial-Temporal Reasoning Through Pretrained Language Models for Video-Grounded Dialogues

Patent №

US 11,487,999

Granted

2022-11-01

Filed 2020

Owner

SALESFORCE.COM, INC.

Lab

AI components

5

ml · nlp · vision · speech · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16860977

A system and method for generating a response in a video grounded dialogue are provided. A video-grounded dialogue neural network language model receives video input and text input. The text input includes a dialogue history between the model and a human user and a current utterance by the user. Encoded video input is generated using video encoding layers. Encoded text input is generated using text encoding layers. The encoded video input and the encoded text input are concatenated in to a single input sequence. A generative pre-trained transformer model generates the response to the current utterance from the singe input sequence.

Machine learningNatural languageVisionSpeechAI hardwareH04N 19/46G06F 40/40G06N 3/045G06N 3/0464G06N 3/047G06N 3/0475G06N 3/049G06N 3/08+10 more

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
AI hardware1.00
Vision0.96
Knowledge representation0.04
Evolutionary computation0.00
Planning0.00

Ownership

SALESFORCE.COM, INC.

assignment · 525280587

Assignors

LE, HUNG, HOI, CHU HONG

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC