Speech synthesis is the artificial production of human speech. A typical\ntext-to-speech system converts a language text into a waveform. There exist\nmany English TTS systems that produce mature, natural, and human-like speech\nsynthesizers. In contrast, other languages, including Arabic, have not been\nconsidered until recently. Existing Arabic speech synthesis solutions are slow,\nof low quality, and the naturalness of synthesized speech is inferior to the\nEnglish synthesizers. They also lack essential speech key factors such as\nintonation, stress, and rhythm. Different works were proposed to solve those\nissues, including the use of concatenative methods such as unit selection or\nparametric methods. However, they required a lot of laborious work and domain\nexpertise. Another reason for such poor performance of Arabic speech\nsynthesizers is the lack of speech corpora, unlike English that has many\npublicly available corpora and audiobooks. This work describes how to generate\nhigh quality, natural, and human-like Arabic speech using an end-to-end neural\ndeep network architecture. This work uses just $\\langle$ text, audio $\\rangle$\npairs with a relatively small amount of recorded audio samples with a total of\n2.41 hours. It illustrates how to use English character embedding despite using\ndiacritic Arabic characters as input and how to preprocess these audio samples\nto achieve the best results.\n
Paper
References (30)
Scroll for more · 18 remaining