Method and System for Building Text-to-Speech Voice from Diverse Recordings

Patent №

US 9,542,927

Granted

2017-01-10

Filed 2014

Owner

GOOGLE INC.

AI components

3

ml · nlp · speech

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

14540088

A method and system is disclosed for building a speech database for a text-to-speech (TTS) synthesis system from multiple speakers recorded under diverse conditions. For a plurality of utterances of a reference speaker, a set of reference-speaker vectors may be extracted, and for each of a plurality of utterances of a colloquial speaker, a respective set of colloquial-speaker vectors may be extracted. A matching procedure, carried out under a transform that compensates for speaker differences, may be used to match each colloquial-speaker vector to a reference-speaker vector. The colloquial-speaker vector may be replaced with the matched reference-speaker vector. The matching-and-replacing can be carried out separately for each set of colloquial-speaker vectors. A conditioned set of speaker vectors can then be constructed by aggregating all the replaced speaker vectors. The condition set of speaker vectors can be used to train the TTS system.

AI classification

Machine learning1.00
Natural language1.00
Speech1.00
AI hardware0.46
Vision0.21
Planning0.13
Evolutionary computation0.00
Knowledge representation0.00

Ownership

GOOGLE INC.

assignment · 341620025

Assignors

AGIOMYRGIANNAKIS, IOANNIS, GUTKIN, ALEXANDER

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC