Multimodal Translation Framework with Cross-Modal Semantic Alignment Using Computer Vision and Language Modeling

The rapid advancement of artificial intelligence has driven the widespread application of cross-modal information processing technologies in language understanding and translation. Addressing the limitations of traditional foreign language translation methods—such as inadequate semantic comprehension in complex text-image contexts and incomplete context modeling—this paper proposes a multimodal foreign language translation framework integrating computer vision and natural language processing. This approach constructs a dual-modal feature extraction pathway for images and text based on a multi-scale residual visual encoder and a pre-trained language model. A dual-path dynamic fusion mechanism achieves deep semantic alignment between images and text. Additionally, a semantic consistency enhancement strategy and a graph-structured supervision module are introduced to constrain cross-modal semantic residuals, thereby improving the stability and accuracy of translation content. Comparative experiments on the Flickr30k and Multi30k datasets demonstrate that the proposed model outperforms existing mainstream models across multiple metrics including BLEU-4, METEOR, ROUGE-L, and CIDEr, achieving a CIDEr score of 118.4. This highlights its robust semantic understanding and generation capabilities. These findings validate the effectiveness of constructing highly coherent, multimodal fusion translation models, providing technical support for intelligent translation systems that integrate images and text.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC