Fast Griffin Lim based Waveform Generation Strategy for Text-to-Speech\n Synthesis

The performance of text-to-speech (TTS) systems heavily depends on\nspectrogram to waveform generation, also known as the speech reconstruction\nphase. The time required for the same is known as synthesis delay. In this\npaper, an approach to reduce speech synthesis delay has been proposed. It aims\nto enhance the TTS systems for real-time applications such as digital\nassistants, mobile phones, embedded devices, etc. The proposed approach applies\nFast Griffin Lim Algorithm (FGLA) instead Griffin Lim algorithm (GLA) as\nvocoder in the speech synthesis phase. GLA and FGLA are both iterative, but the\nconvergence rate of FGLA is faster than GLA. The proposed approach is tested on\nLJSpeech, Blizzard and Tatoeba datasets and the results for FGLA are compared\nagainst GLA and neural Generative Adversarial Network (GAN) based vocoder. The\nperformance is evaluated based on synthesis delay and speech quality. A 36.58%\nreduction in speech synthesis delay has been observed. The quality of the\noutput speech has improved, which is advocated by higher Mean opinion scores\n(MOS) and faster convergence with FGLA as opposed to GLA.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC