This research-based project is about a new way to put feelings into computer-generated speech. It has two stages: text emotion detection and emotional speech synthesis. In the former part, labeled text data is used to build models to spot the emotions. This model learns to change pitch, volume, and rhythm based on the given emotion. During the latter phase, the text is changed into neutral speech using text-to-speech (TTS) methods. The model is used for practice to find the emotion in the text, tagging the text with that emotion. Also, a deep-learning model known as a tacotron is employed to make an emotional-sounding speech. The model perfectly uses machine learning strategies to blend feelings into speech. This opens up possibilities for turning plain text into speech with feelings of varied emotions. The proposed system makes a novice and important contribution to natural language processing for emotion identification and deep learning for emotional speech synthesis.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex