Multimodal Translation Fusion Model based on Attention Mechanism Algorithm

With the growing complexity of cross-language communication, traditional unimodal translation models struggle to handle intricate contextual understanding and multimodal interactions. This paper proposes a multimodal translation fusion model leveraging the self-attention mechanism within the Transformer architecture to address these challenges. We first pre-process text, image, and speech data independently using modality-specific encoders. Next, a cross-modal attention layer dynamically aligns and integrates features by computing inter-modal correlations, prioritizing contextually salient information. Finally, a multi-layer Transformer refines fused representations for translation generation. Experiments demonstrate significant improvements, achieving BLEU scores of 0.7325 (2.5% higher than unimodal baselines) and ROUGE/METEOR gains of 28.7% and 18%, respectively. The model excels in Latin language pairs (e.g., EN-FR/ES) but shows limitations in non-Latin languages (e.g., EN-JA), highlighting its context-aware adaptability while underscoring the need for structural optimization in diverse linguistic systems. This work validates the efficacy of attention-driven multimodal fusion in enhancing translation coherence and cross-lingual robustness.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC