Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech

In this paper, we introduce Kathaka, a model trained with a novel two-stage\ntraining process for neural speech synthesis with contextually appropriate\nprosody. In Stage I, we learn a prosodic distribution at the sentence level\nfrom mel-spectrograms available during training. In Stage II, we propose a\nnovel method to sample from this learnt prosodic distribution using the\ncontextual information available in text. To do this, we use BERT on text, and\ngraph-attention networks on parse trees extracted from text. We show a\nstatistically significant relative improvement of $13.2\\%$ in naturalness over\na strong baseline when compared to recordings. We also conduct an ablation\nstudy on variations of our sampling technique, and show a statistically\nsignificant improvement over the baseline in each case.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC