This paper proposes a visual-text to speech (vTTS) method, a method for synthesizing speech directly from visual text (i.e., text as an image). vTTS can use visual features in visual text that should be important for speech synthesis such as emphasis and radicals (components in Chinese characters), but they are not available in conventional TTS using discrete symbols as the input. The proposed vTTS method extracts visual features with a convolutional neural network and then generates acoustic features with a non-autoregressive model in an end-to-end manner. Experimental results show that 1) the vTTS method is capable of generating speech with naturalness comparable to or better than a conventional TTS, 2) it can transfer emphasis and emotion attributes in visual text to speech without additional labels and architectures, and 3) it can synthesize more natural and intelligible speech from unseen and rare characters than conventional TTS.
Paper
References (36)
Scroll for more · 24 remaining