In the context of the Fourth Industrial Revolution, which is strongly driving comprehensive digital transformation, building natural communication methods between humans and machines through speech has become a central task in the field of artificial intelligence. This comprehensive study systematizes the theoretical and practical foundations of text-to-speech (TTS) technology specifically for the Vietnamese language. We applied a comprehensive research method, including in-depth analysis of more than 40 reputable scientific sources, to clarify the shift towards modern deep neural networks. The research focused on the operating mechanisms of architectures such as FastPitch and machine learning techniques that do not require sample data to address the problems of naturalness and emotional expressiveness in native speech. The results confirm that deep integration of phonetic features enhances system performance and significantly contributes to the preservation of national digital cultural heritage.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex