Impacts of input linguistic feature representation on Japanese end-to-end speech synthesis

We investigate the impact of input linguistic feature representation on Japanese end-to-end speech synthesis. An end-to-end speech synthesis system, which directly generates natural speech from text, has recently been proposed. The English end-to-end system Tacotron 2 achieves sound quality close to that of natural speech. However, unlike alphabetic language that use stress accent, such as English and Spanish, it is difficult to achieve end-to-end speech synthesis with other non-alphabetic languages (e.g., Japanese and Chinese, which use pitch accent and tone, respectively, and use ideograms). We investigated the units of an input sequence, contexts, pause insertion, vowel de-voicing, and pronunciation of particles for Japanese end-to-end speech synthesis. Experimental results indicate improvement in the naturalness of the synthesized speech using high or low accents. The results also indicate that the accent-phrase information can help to predict pause insertion, and an end-to-end text-to-speech model may be able to change the pronunciation for devoiced vowels and particles.

Paper

Full text

PDF

Impacts of input linguistic feature representation on Japanese end-to-end speech synthesis

Semantic Scholar · Linguistics · 2019

Abstract

We investigate the impact of input linguistic feature representation on Japanese end-to-end speech synthesis. An end-to-end speech synthesis system, which directly generates natural speech from text, has recently been proposed. The English end-to-end system Tacotron 2 achieves sound quality close to that of natural speech. However, unlike alphabetic language that use stress accent, such as English and Spanish, it is difficult to achieve end-to-end speech synthesis with other non-alphabetic languages (e.g., Japanese and Chinese, which use pitch accent and tone, respectively, and use ideograms). We investigated the units of an input sequence, contexts, pause insertion, vowel de-voicing, and pronunciation of particles for Japanese end-to-end speech synthesis. Experimental results indicate improvement in the naturalness of the synthesized speech using high or low accents. The results also indicate that the accent-phrase information can help to predict pause insertion, and an end-to-end text-to-speech model may be able to change the pronunciation for devoiced vowels and particles.

Similar papers

© 2026 NYSGPT2525 LLC