Representation Learning for Conversational Data using Discourse Mutual Information Maximization

Although many pretrained models exist for text or images, there have been\nrelatively fewer attempts to train representations specifically for dialog\nunderstanding. Prior works usually relied on finetuned representations based on\ngeneric text representation models like BERT or GPT-2. But such language\nmodeling pretraining objectives do not take the structural information of\nconversational text into consideration. Although generative dialog models can\nlearn structural features too, we argue that the structure-unaware word-by-word\ngeneration is not suitable for effective conversation modeling. We empirically\ndemonstrate that such representations do not perform consistently across\nvarious dialog understanding tasks. Hence, we propose a structure-aware Mutual\nInformation based loss-function DMI (Discourse Mutual Information) for training\ndialog-representation models, that additionally captures the inherent\nuncertainty in response prediction. Extensive evaluation on nine diverse dialog\nmodeling tasks shows that our proposed DMI-based models outperform strong\nbaselines by significant margins.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC