DEAM: Dialogue Coherence Evaluation using AMR-based Semantic Manipulations

Automatic evaluation metrics are essential for the rapid development of\nopen-domain dialogue systems as they facilitate hyper-parameter tuning and\ncomparison between models. Although recently proposed trainable\nconversation-level metrics have shown encouraging results, the quality of the\nmetrics is strongly dependent on the quality of training data. Prior works\nmainly resort to heuristic text-level manipulations (e.g. utterances shuffling)\nto bootstrap incoherent conversations (negative examples) from coherent\ndialogues (positive examples). Such approaches are insufficient to\nappropriately reflect the incoherence that occurs in interactions between\nadvanced dialogue models and humans. To tackle this problem, we propose DEAM, a\nDialogue coherence Evaluation metric that relies on Abstract Meaning\nRepresentation (AMR) to apply semantic-level Manipulations for incoherent\n(negative) data generation. AMRs naturally facilitate the injection of various\ntypes of incoherence sources, such as coreference inconsistency, irrelevancy,\ncontradictions, and decrease engagement, at the semantic level, thus resulting\nin more natural incoherent samples. Our experiments show that DEAM achieves\nhigher correlations with human judgments compared to baseline methods on\nseveral dialog datasets by significant margins. We also show that DEAM can\ndistinguish between coherent and incoherent dialogues generated by baseline\nmanipulations, whereas those baseline models cannot detect incoherent examples\ngenerated by DEAM. Our results demonstrate the potential of AMR-based semantic\nmanipulations for natural negative example generation.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC