A Coordinated Representation Learning Enhanced Multimodal Machine Translation Approach with Multi-Attention
In recent years, the application of machine translation has become more and more widely. Currently, the neural multimodal translation models have made attractive progress, which combines images into deep learning networks, such as Transformer and RNN. When considering images in translation models, they directly apply gate structure or image attention to introduce image feature to enhance the translation effect. We argue that it may mismatch the text and image features since they are in different semantic space. In this paper, we propose a coordinated representation learning enhanced multimodal machine translation approach with multimodal attention. Our approach accepts the text data and its relevant image data as the input. The image features are fed into the decoder side of the basic Transformer model. Moreover, the Coordinated Representation Learning is utilized to map the different text and image modal features into their semantic representations. The mapped representations are linearly related in a shared semantic space. Finally, the sum of the image and text representations, called Coordinated Visual-Semantic Representation (CVSR), will be sent to a Multimodal Attention Layer (MAL) in our Transformer based translation approach. Experimental results show that our approach achieves the state-of-art performance on the public Multi30k dataset.
Paper
Full text
A Coordinated Representation Learning Enhanced Multimodal Machine Translation Approach with Multi-Attention
OpenAlex · Natural Language Processing Techniques · 2020
Abstract
In recent years, the application of machine translation has become more and more widely. Currently, the neural multimodal translation models have made attractive progress, which combines images into deep learning networks, such as Transformer and RNN. When considering images in translation models, they directly apply gate structure or image attention to introduce image feature to enhance the translation effect. We argue that it may mismatch the text and image features since they are in different semantic space. In this paper, we propose a coordinated representation learning enhanced multimodal machine translation approach with multimodal attention. Our approach accepts the text data and its relevant image data as the input. The image features are fed into the decoder side of the basic Transformer model. Moreover, the Coordinated Representation Learning is utilized to map the different text and image modal features into their semantic representations. The mapped representations are linearly related in a shared semantic space. Finally, the sum of the image and text representations, called Coordinated Visual-Semantic Representation (CVSR), will be sent to a Multimodal Attention Layer (MAL) in our Transformer based translation approach. Experimental results show that our approach achieves the state-of-art performance on the public Multi30k dataset.