What Helps Transformers Recognize Conversational Structure? Importance of Context, Punctuation, and Labels in Dialog Act Recognition

Dialog acts can be interpreted as the atomic units of a conversation, more\nfine-grained than utterances, characterized by a specific communicative\nfunction. The ability to structure a conversational transcript as a sequence of\ndialog acts -- dialog act recognition, including the segmentation -- is\ncritical for understanding dialog. We apply two pre-trained transformer models,\nXLNet and Longformer, to this task in English and achieve strong results on\nSwitchboard Dialog Act and Meeting Recorder Dialog Act corpora with dialog act\nsegmentation error rates (DSER) of 8.4% and 14.2%. To understand the key\nfactors affecting dialog act recognition, we perform a comparative analysis of\nmodels trained under different conditions. We find that the inclusion of a\nbroader conversational context helps disambiguate many dialog act classes,\nespecially those infrequent in the training data. The presence of punctuation\nin the transcripts has a massive effect on the models' performance, and a\ndetailed analysis reveals specific segmentation patterns observed in its\nabsence. Finally, we find that the label set specificity does not affect dialog\nact segmentation performance. These findings have significant practical\nimplications for spoken language understanding applications that depend heavily\non a good-quality segmentation being available.\n

Paper

References (61)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC