Leverage Unlabeled Data for Abstractive Speech Summarization with Self-Supervised Learning and Back-Summarization

Supervised approaches for Neural Abstractive Summarization require large\nannotated corpora that are costly to build. We present a French meeting\nsummarization task where reports are predicted based on the automatic\ntranscription of the meeting audio recordings. In order to build a corpus for\nthis task, it is necessary to obtain the (automatic or manual) transcription of\neach meeting, and then to segment and align it with the corresponding manual\nreport to produce training examples suitable for training. On the other hand,\nwe have access to a very large amount of unaligned data, in particular reports\nwithout corresponding transcription. Reports are professionally written and\nwell formatted making pre-processing straightforward. In this context, we study\nhow to take advantage of this massive amount of unaligned data using two\napproaches (i) self-supervised pre-training using a target-side denoising\nencoder-decoder model; (ii) back-summarization i.e. reversing the summarization\nprocess by learning to predict the transcription given the report, in order to\nalign single reports with generated transcription, and use this synthetic\ndataset for further training. We report large improvements compared to the\nprevious baseline (trained on aligned data only) for both approaches on two\nevaluation sets. Moreover, combining the two gives even better results,\noutperforming the baseline by a large margin of +6 ROUGE-1 and ROUGE-L and +5\nROUGE-2 on two evaluation sets\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC