TRANSLATING VIDEO TO LANGUAGE USING ADAPTIVE SPATIOTEMPORAL CONVOLUTION FEATURE REPRESENTATION WITH DYNAMIC ABSTRACTION

Patent №

US 10,366,292

Granted

2019-07-30

Filed 2017

Owner

NEC LABORATORIES AMERICA, INC.

Lab

AI components

6

ml · nlp · vision · speech · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

15794758

A system is provided for video captioning. The system includes a processor. The processor is configured to apply a three-dimensional Convolutional Neural Network (C3D) to image frames of a video sequence to obtain, for the video sequence, (i) intermediate feature representations across L convolutional layers and (ii) top-layer features. The processor is further configured to produce a first word of an output caption for the video sequence by applying the top-layer features to a Long Short Term Memory (LSTM). The processor is further configured to produce subsequent words of the output caption by (i) dynamically performing spatiotemporal attention and layer attention using the intermediate feature representations to form a context vector, and (ii) applying the LSTM to the context vector, a previous word of the output caption, and a hidden state of the LSTM. The system further includes a display device for displaying the output caption to a user.

Machine learningNatural languageVisionSpeechKnowledge representationAI hardwareH04N 7/183G06F 18/2148G06F 18/2415G06N 3/044G06N 3/0442G06N 3/045G06N 3/0455G06N 3/0464+15 more

AI classification

Machine learning1.00
Natural language1.00
Speech1.00
Vision1.00
AI hardware1.00
Knowledge representation0.96
Planning0.16
Evolutionary computation0.00

Ownership

NEC LABORATORIES AMERICA, INC.

assignment · 439610473

Assignors

MIN, RENQIANG, PU, YUNCHEN

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC