Multiway Attention for Neural Machine Translation

Neural machine translation (NMT) with source side attention has achieved remarkable performance. Nevertheless, all existing attention mechanisms employ only one attention function. However, several different attention functions have been proposed. They have different mechanisms and capture different information of the sentences, thus a single attention function does not perform well. In the paper, we propose the multiway attention neural machine translation model (MA-NMT) which employs multiple attention functions in the attention mechanism to calculate the weight of each source word when predicting the next target word. Specially, we design three attention functions to get the contextual information. Then we combine the features from all attention function to obtain the final semantic representation. The results of experiments on the English-German translation task demonstrate that the proposed MA-NMT improves the performance than the baseline NMT models.

Paper

Full text

PDF

Multiway Attention for Neural Machine Translation

Semantic Scholar · Computer Science · 2019

Abstract

Neural machine translation (NMT) with source side attention has achieved remarkable performance. Nevertheless, all existing attention mechanisms employ only one attention function. However, several different attention functions have been proposed. They have different mechanisms and capture different information of the sentences, thus a single attention function does not perform well. In the paper, we propose the multiway attention neural machine translation model (MA-NMT) which employs multiple attention functions in the attention mechanism to calculate the weight of each source word when predicting the next target word. Specially, we design three attention functions to get the contextual information. Then we combine the features from all attention function to obtain the final semantic representation. The results of experiments on the English-German translation task demonstrate that the proposed MA-NMT improves the performance than the baseline NMT models.

Similar papers

© 2026 NYSGPT2525 LLC