MULTIMODAL DATA PROCESSING

Patent №

US 12,675,987

Granted

2026-07-07

Filed 2023

Owner

Beijing Youzhuju Network Technology Co., Ltd.

Lab

AI components

0

Assignment

None on record

Dataset

AIPD

Application

18393238

Embodiments of the present disclosure provide a solution for multimodal data processing. A method comprises: obtaining image data and text data; and extracting a target visual feature of image data and a target textual feature of text data using a feature extraction model. The feature extraction model comprises alternatively deployed cross-modal encoding parts and visual encoding parts. The extracting comprises: performing, using a first cross-modal encoding part of the feature extraction model, cross-modal feature encoding on a first intermediate visual feature of the image data and a first intermediate textual feature of the text data, to obtain a second intermediate visual feature and a second intermediate textual feature; performing, using a first visual encoding part of the feature extraction model, visual modal feature encoding on the second intermediate visual feature, to obtain a third intermediate visual feature.

G06V 10/82G06V 10/467

Ownership

Beijing Youzhuju Network Technology Co., Ltd.

From the same owner

© 2026 NYSGPT2525 LLC