Enhancing Speech Intelligibility in Text-To-Speech Synthesis using Speaking Style Conversion

The increased adoption of digital assistants makes text-to-speech (TTS)\nsynthesis systems an indispensable feature of modern mobile devices. It is\nhence desirable to build a system capable of generating highly intelligible\nspeech in the presence of noise. Past studies have investigated style\nconversion in TTS synthesis, yet degraded synthesized quality often leads to\nworse intelligibility. To overcome such limitations, we proposed a novel\ntransfer learning approach using Tacotron and WaveRNN based TTS synthesis. The\nproposed speech system exploits two modification strategies: (a) Lombard\nspeaking style data and (b) Spectral Shaping and Dynamic Range Compression\n(SSDRC) which has been shown to provide high intelligibility gains by\nredistributing the signal energy on the time-frequency domain. We refer to this\nextension as Lombard-SSDRC TTS system. Intelligibility enhancement as\nquantified by the Intelligibility in Bits (SIIB-Gauss) measure shows that the\nproposed Lombard-SSDRC TTS system shows significant relative improvement\nbetween 110% and 130% in speech-shaped noise (SSN), and 47% to 140% in\ncompeting-speaker noise (CSN) against the state-of-the-art TTS approach.\nAdditional subjective evaluation shows that Lombard-SSDRC TTS successfully\nincreases the speech intelligibility with relative improvement of 455% for SSN\nand 104% for CSN in median keyword correction rate compared to the baseline TTS\nmethod.\n

Paper

References (28)

Scroll for more · 16 remaining

Similar papers

© 2026 NYSGPT2525 LLC