Recent advances in language and vision push forward the research of\ncaptioning a single image to describing visual differences between image pairs.\nSuppose there are two images, I_1 and I_2, and the task is to generate a\ndescription W_{1,2} comparing them, existing methods directly model { I_1, I_2\n} -> W_{1,2} mapping without the semantic understanding of individuals. In this\npaper, we introduce a Learning-to-Compare (L2C) model, which learns to\nunderstand the semantic structures of these two images and compare them while\nlearning to describe each one. We demonstrate that L2C benefits from a\ncomparison between explicit semantic representations and single-image captions,\nand generalizes better on the new testing image pairs. It outperforms the\nbaseline on both automatic evaluation and human evaluation for the\nBirds-to-Words dataset.\n
Paper
References (20)
Scroll for more · 8 remaining