An objective evaluation of the effects of recording conditions and\n speaker characteristics in multi-speaker deep neural speech synthesis
Multi-speaker spoken datasets enable the creation of text-to-speech synthesis\n(TTS) systems which can output several voice identities. The multi-speaker\n(MSPK) scenario also enables the use of fewer training samples per speaker.\nHowever, in the resulting acoustic model, not all speakers exhibit the same\nsynthetic quality, and some of the voice identities cannot be used at all.\n In this paper we evaluate the influence of the recording conditions, speaker\ngender, and speaker particularities over the quality of the synthesised output\nof a deep neural TTS architecture, namely Tacotron2. The evaluation is possible\ndue to the use of a large Romanian parallel spoken corpus containing over 81\nhours of data. Within this setup, we also evaluate the influence of different\ntypes of text representations: orthographic, phonetic, and phonetic extended\nwith syllable boundaries and lexical stress markings.\n We evaluate the results of the MSPK system using the objective measures of\nequal error rate (EER) and word error rate (WER), and also look into the\ndistances between natural and synthesised t-SNE projections of the embeddings\ncomputed by an accurate speaker verification network. The results show that\nthere is indeed a large correlation between the recording conditions and the\nspeaker's synthetic voice quality. The speaker gender does not influence the\noutput, and that extending the input text representation with syllable\nboundaries and lexical stress information does not equally enhance the\ngenerated audio across all speaker identities. The visualisation of the t-SNE\nprojections of the natural and synthesised speaker embeddings show that the\nacoustic model shifts some of the speakers' neural representation, but not all\nof them. As a result, these speakers have lower performances of the output\nspeech.\n