A Corpus of Controlled Opinionated and Knowledgeable Movie Discussions for Training Neural Conversation Models

Fully data driven Chatbots for non-goal oriented dialogues are known to\nsuffer from inconsistent behaviour across their turns, stemming from a general\ndifficulty in controlling parameters like their assumed background personality\nand knowledge of facts. One reason for this is the relative lack of labeled\ndata from which personality consistency and fact usage could be learned\ntogether with dialogue behaviour. To address this, we introduce a new labeled\ndialogue dataset in the domain of movie discussions, where every dialogue is\nbased on pre-specified facts and opinions. We thoroughly validate the collected\ndialogue for adherence of the participants to their given fact and opinion\nprofile, and find that the general quality in this respect is high. This\nprocess also gives us an additional layer of annotation that is potentially\nuseful for training models. We introduce as a baseline an end-to-end trained\nself-attention decoder model trained on this data and show that it is able to\ngenerate opinionated responses that are judged to be natural and knowledgeable\nand show attentiveness.\n

Paper

References (25)

Scroll for more · 13 remaining

Similar papers

© 2026 NYSGPT2525 LLC