Using IPA-Based Tacotron for Data Efficient Cross-Lingual Speaker Adaptation and Pronunciation Enhancement
Recent neural Text-to-Speech (TTS) models have been shown to perform very\nwell when enough data is available. However, fine-tuning them for new speakers\nor languages is not straightforward in a low-resource setup. In this paper, we\nshow that by applying minor modifications to a Tacotron model, one can transfer\nan existing TTS model for new speakers from the same or a different language\nusing only 20 minutes of data. For this purpose, we first introduce a base\nmulti-lingual Tacotron with language-agnostic input, then demonstrate how\ntransfer learning is done for different scenarios of speaker adaptation without\nexploiting any pre-trained speaker encoder or code-switching technique. We\nevaluate the transferred model in both subjective and objective ways.\n