Multimodal translation has to cope with continuous challenges such as feature misalignment, temporal inconsistencies, and limited handling of out-of-vocabulary words. In an attempt to overcome these drawbacks, we propose a Hierarchical Attention Temporal Sequence model (HATSeq-LSTM), wherein textual, visual, and audio information is integrated through dual-level attention and dynamic fusion. To extract information, the model uses TF-IDF and a bidirectional GRU with attention for text, VGG-16 for image features, and MFCCs for audio signals, followed by an encoder-decoder LSTM to generate the translation. Experimental results on benchmark multimodal translation datasets demonstrate that HATSeq-LSTM outperforms existing methods, achieving BLEU, ROUGE, and METEOR scores consistently above 90%, highlighting its effectiveness in cross-modal integration and translation quality.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex