When speakers describe an image, they tend to look at objects before\nmentioning them. In this paper, we investigate such sequential cross-modal\nalignment by modelling the image description generation process\ncomputationally. We take as our starting point a state-of-the-art image\ncaptioning system and develop several model variants that exploit information\nfrom human gaze patterns recorded during language production. In particular, we\npropose the first approach to image description generation where visual\nprocessing is modelled $\\textit{sequentially}$. Our experiments and analyses\nconfirm that better descriptions can be obtained by exploiting gaze-driven\nattention and shed light on human cognitive processes by comparing different\nways of aligning the gaze modality with language production. We find that\nprocessing gaze data sequentially leads to descriptions that are better aligned\nto those produced by speakers, more diverse, and more natural${-}$particularly\nwhen gaze is encoded with a dedicated recurrent component.\n
Paper
References (53)
Scroll for more · 38 remaining