What Helps Transformers Recognize Conversational Structure? Importance of Context, Punctuation, and Labels in Dialog Act Recognition
Dialog acts can be interpreted as the atomic units of a conversation, more\nfine-grained than utterances, characterized by a specific communicative\nfunction. The ability to structure a conversational transcript as a sequence of\ndialog acts -- dialog act recognition, including the segmentation -- is\ncritical for understanding dialog. We apply two pre-trained transformer models,\nXLNet and Longformer, to this task in English and achieve strong results on\nSwitchboard Dialog Act and Meeting Recorder Dialog Act corpora with dialog act\nsegmentation error rates (DSER) of 8.4% and 14.2%. To understand the key\nfactors affecting dialog act recognition, we perform a comparative analysis of\nmodels trained under different conditions. We find that the inclusion of a\nbroader conversational context helps disambiguate many dialog act classes,\nespecially those infrequent in the training data. The presence of punctuation\nin the transcripts has a massive effect on the models' performance, and a\ndetailed analysis reveals specific segmentation patterns observed in its\nabsence. Finally, we find that the label set specificity does not affect dialog\nact segmentation performance. These findings have significant practical\nimplications for spoken language understanding applications that depend heavily\non a good-quality segmentation being available.\n
Paper
References (61)
Scroll for more · 38 remaining