Patent №
US 11,281,945
Granted
2022-03-22
Filed 2021
Owner
INSTITUTE OF AUTOMATION, CHINESE ACADEMY OF SCIENCES
Lab
—
AI components
5
ml · nlp · vision · speech · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
17468994
A multimodal dimensional emotion recognition method includes: acquiring a frame-level audio feature, a frame-level video feature, and a frame-level text feature from an audio, a video, and a corresponding text of a sample to be tested; performing temporal contextual modeling on the frame-level audio feature, the frame-level video feature, and the frame-level text feature respectively by using a temporal convolutional network to obtain a contextual audio feature, a contextual video feature, and a contextual text feature; performing weighted fusion on these three features by using a gated attention mechanism to obtain a multimodal feature; splicing the multimodal feature and these three features together to obtain a spliced feature, and then performing further temporal contextual modeling on the spliced feature by using a temporal convolutional network to obtain a contextual spliced feature; and performing regression prediction on the contextual spliced feature to obtain a final dimensional emotion prediction result.
AI classification
Ownership
INSTITUTE OF AUTOMATION, CHINESE ACADEMY OF SCIENCES
assignment · 574110646
Assignors
TAO, JIANHUA, SUN, LICAI, LIU, BIN, LIAN, ZHENG
On an employer assignment, the assignors are typically the inventors.