Reinforce-Aligner: Reinforcement Alignment Search for Robust End-to-End\n Text-to-Speech

Text-to-speech (TTS) synthesis is the process of producing synthesized speech\nfrom text or phoneme input. Traditional TTS models contain multiple processing\nsteps and require external aligners, which provide attention alignments of\nphoneme-to-frame sequences. As the complexity increases and efficiency\ndecreases with every additional step, there is expanding demand in modern\nsynthesis pipelines for end-to-end TTS with efficient internal aligners. In\nthis work, we propose an end-to-end text-to-waveform network with a novel\nreinforcement learning based duration search method. Our proposed generator is\nfeed-forward and the aligner trains the agent to make optimal duration\npredictions by receiving active feedback from actions taken to maximize\ncumulative reward. We demonstrate accurate alignments of phoneme-to-frame\nsequence generated from trained agents enhance fidelity and naturalness of\nsynthesized audio. Experimental results also show the superiority of our\nproposed model compared to other state-of-the-art TTS models with internal and\nexternal aligners.\n

Paper

References (36)

Scroll for more · 24 remaining

Similar papers

© 2026 NYSGPT2525 LLC