Exploring Solutions for Text-to-Speech Synthesis of Low-Resource Languages

A text-to-speech synthesis system is expected to convert any given text to highly intelligible and natural speech that sounds as human-like as possible. It finds its place in a variety of applications, like virtual assistants, speech-assistive aids, and speech-to-speech translation systems. In recent years several deep learning-based methods have been proposed for text-to-speech synthesis, which result in synthetic speech that is intelligible and very natural. However, these methods require a large amount of speech data, in the order of tens of hours, preferably recorded from a single speaker and in a quiet/studio environment. While such a large amount of data might be easily available for certain languages, like English, which are widely spoken and used globally, this is not the case with several other languages. In this regard, the current work explores state-of-the-art text-to-speech synthesis methods and strategies for handling low-resource languages. The paper ultimately aims to provide insight into developing text-to-speech systems for Indian languages and hence reviews the work done so far on Indian languages as well. Since India is a multilingual country where several languages co-exist, a multilingual synthesizer would be preferable. Further, one recurring suggestion to handle data scarcity is to make use of data from multiple languages too. Therefore, the current work focuses on exploring multilingual text-to-speech synthesis for low-resource languages.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC