Dialogue participants often refer to entities or situations repeatedly within\na conversation, which contributes to its cohesiveness. Subsequent references\nexploit the common ground accumulated by the interlocutors and hence have\nseveral interesting properties, namely, they tend to be shorter and reuse\nexpressions that were effective in previous mentions. In this paper, we tackle\nthe generation of first and subsequent references in visually grounded\ndialogue. We propose a generation model that produces referring utterances\ngrounded in both the visual and the conversational context. To assess the\nreferring effectiveness of its output, we also implement a reference resolution\nsystem. Our experiments and analyses show that the model produces better, more\neffective referring utterances than a model not grounded in the dialogue\ncontext, and generates subsequent references that exhibit linguistic patterns\nakin to humans.\n
Paper
References (68)
Scroll for more · 38 remaining