Joint prediction of truecasing and punctuation for conversational speech in low-resource scenarios

Capitalization and punctuation are important cues for comprehending written\ntexts and conversational transcripts. Yet, many ASR systems do not produce\npunctuated and case-formatted speech transcripts. We propose to use a\nmulti-task system that can exploit the relations between casing and punctuation\nto improve their prediction performance. Whereas text data for predicting\npunctuation and truecasing is seemingly abundant, we argue that written text\nresources are inadequate as training data for conversational models. We\nquantify the mismatch between written and conversational text domains by\ncomparing the joint distributions of punctuation and word cases, and by testing\nour model cross-domain. Further, we show that by training the model in the\nwritten text domain and then transfer learning to conversations, we can achieve\nreasonable performance with less data.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC