Multimodal Semi-supervised Learning Framework for Punctuation Prediction in Conversational Speech
In this work, we explore a multimodal semi-supervised learning approach for\npunctuation prediction by learning representations from large amounts of\nunlabelled audio and text data. Conventional approaches in speech processing\ntypically use forced alignment to encoder per frame acoustic features to word\nlevel features and perform multimodal fusion of the resulting acoustic and\nlexical representations. As an alternative, we explore attention based\nmultimodal fusion and compare its performance with forced alignment based\nfusion. Experiments conducted on the Fisher corpus show that our proposed\napproach achieves ~6-9% and ~3-4% absolute improvement (F1 score) over the\nbaseline BLSTM model on reference transcripts and ASR outputs respectively. We\nfurther improve the model robustness to ASR errors by performing data\naugmentation with N-best lists which achieves up to an additional ~2-6%\nimprovement on ASR outputs. We also demonstrate the effectiveness of\nsemi-supervised learning approach by performing ablation study on various sizes\nof the corpus. When trained on 1 hour of speech and text data, the proposed\nmodel achieved ~9-18% absolute improvement over baseline model.\n