Controllable Sequence-To-Sequence Neural TTS with LPCNET Backend for\n Real-time Speech Synthesis on CPU

State-of-the-art sequence-to-sequence acoustic networks, that convert a\nphonetic sequence to a sequence of spectral features with no explicit prosody\nprediction, generate speech with close to natural quality, when cascaded with\nneural vocoders, such as Wavenet. However, the combined system is typically too\nheavy for real-time speech synthesis on a CPU. In this work we present a\nsequence-to-sequence acoustic network combined with lightweight LPCNet neural\nvocoder, designed for real-time speech synthesis on a CPU. In addition, the\nsystem allows sentence-level pace and expressivity control at inference time.\nWe demonstrate that the proposed system can synthesize high quality 22 kHz\nspeech in real-time on a general-purpose CPU. In terms of MOS score degradation\nrelative to PCM, the system attained as low as 6.1-6.5% for quality and 6.3-\n7.0% for expressiveness, reaching equivalent or better quality when compared to\na similar system with a Wavenet vocoder backend.\n

Paper

References (14)

Scroll for more · 2 remaining

Similar papers

© 2026 NYSGPT2525 LLC