In recent years, statistical parametric speech synthesis (SPSS) systems have\nbeen widely utilized in many interactive speech-based systems (e.g.~Amazon's\nAlexa, Bose's headphones). To select a suitable SPSS system, both speech\nquality and performance efficiency (e.g.~decoding time) must be taken into\naccount. In the paper, we compared four popular Vietnamese SPSS techniques\nusing: 1) hidden Markov models (HMM), 2) deep neural networks (DNN), 3)\ngenerative adversarial networks (GAN), and 4) end-to-end (E2E) architectures,\nwhich consists of Tacontron~2 and WaveGlow vocoder in terms of speech quality\nand performance efficiency. We showed that the E2E systems accomplished the\nbest quality, but required the power of GPU to achieve real-time performance.\nWe also showed that the HMM-based system had inferior speech quality, but it\nwas the most efficient system. Surprisingly, the E2E systems were more\nefficient than the DNN and GAN in inference on GPU. Surprisingly, the GAN-based\nsystem did not outperform the DNN in term of quality.\n
Paper
References (45)
Scroll for more · 33 remaining