Multi-Attention Generative Adversarial Network for image captioning

Abstract Recently, it has been shown that generative-adversarial-nets (GANs) can be directly utilized as an extension of traditional reinforcement-learning in image captioning tasks. However, the GANs-based methods generate captions as a function of only local points in the feature map without capturing non-local information. In this paper, a Multi-Attention mechanism is first proposed by utilizing both of the local and non-local evidence for more effective feature representation and reasoning in image captioning. Based on the mechanism, a Multi-Attention Generative Adversarial Image Captioning Network (MAGAN) is also proposed which contains a Multi-Attention generator and a Multi-Attention discriminator. The proposed generator is designed to generate more accurate sentences, while the proposed discriminator is employed to determine whether generated sentences are human described or machine generated. Extensive experiments are conducted to validate the proposed framework on MSCOCO benchmark dataset, and it achieves very competitive results evaluated by the evaluation server of MS COCO captioning challenge.

Paper

Full text

PDF

Multi-Attention Generative Adversarial Network for image captioning

Semantic Scholar · Computer Science · 2020

Abstract

Abstract Recently, it has been shown that generative-adversarial-nets (GANs) can be directly utilized as an extension of traditional reinforcement-learning in image captioning tasks. However, the GANs-based methods generate captions as a function of only local points in the feature map without capturing non-local information. In this paper, a Multi-Attention mechanism is first proposed by utilizing both of the local and non-local evidence for more effective feature representation and reasoning in image captioning. Based on the mechanism, a Multi-Attention Generative Adversarial Image Captioning Network (MAGAN) is also proposed which contains a Multi-Attention generator and a Multi-Attention discriminator. The proposed generator is designed to generate more accurate sentences, while the proposed discriminator is employed to determine whether generated sentences are human described or machine generated. Extensive experiments are conducted to validate the proposed framework on MSCOCO benchmark dataset, and it achieves very competitive results evaluated by the evaluation server of MS COCO captioning challenge.

Similar papers

© 2026 NYSGPT2525 LLC