StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization

In recent years, neural vocoders have surpassed classical speech generation\napproaches in naturalness and perceptual quality of the synthesized speech.\nComputationally heavy models like WaveNet and WaveGlow achieve best results,\nwhile lightweight GAN models, e.g. MelGAN and Parallel WaveGAN, remain inferior\nin terms of perceptual quality. We therefore propose StyleMelGAN, a lightweight\nneural vocoder allowing synthesis of high-fidelity speech with low\ncomputational complexity. StyleMelGAN employs temporal adaptive normalization\nto style a low-dimensional noise vector with the acoustic features of the\ntarget speech. For efficient training, multiple random-window discriminators\nadversarially evaluate the speech signal analyzed by a filter bank, with\nregularization provided by a multi-scale spectral reconstruction loss. The\nhighly parallelizable speech generation is several times faster than real-time\non CPUs and GPUs. MUSHRA and P.800 listening tests show that StyleMelGAN\noutperforms prior neural vocoders in copy-synthesis and Text-to-Speech\nscenarios.\n

Paper

References (29)

Scroll for more · 17 remaining

Similar papers

© 2026 NYSGPT2525 LLC