Patent №
US 8,315,872
Granted
2012-11-20
Filed 2011
Owner
AT&T CORP.
Lab
—
AI components
2
nlp · speech
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
13306157
A speech synthesis system can select recorded speech fragments, or acoustic units, from a very large database of acoustic units to produce artificial speech. The selected acoustic units are chosen to minimize a combination of target and concatenation costs for a given sentence. However, as concatenation costs, which are measures of the mismatch between sequential pairs of acoustic units, are expensive to compute, processing can be greatly reduced by pre-computing and caching the concatenation costs. Unfortunately, the number of possible sequential pairs of acoustic units makes such caching prohibitive. However, statistical experiments reveal that while about 85% of the acoustic units are typically used in common speech, less than 1% of the possible sequential pairs of acoustic units occur in practice. A method for constructing an efficient concatenation cost database is provided by synthesizing a large body of speech, identifying the acoustic unit sequential pairs generated and their respective concatenation costs, and storing those concatenation costs likely to occur. By constructing a concatenation cost database in this faction, the processing power required at run-time is greatly reduced with negligible effect on speech quality.
AI classification
Ownership
AT&T CORP.
assignment · 272920981
Assignors
BEUTNAGEL, MARK CHARLES, MOHRI, MEHRYAR, RILEY, MICHAEL DENNIS
On an employer assignment, the assignors are typically the inventors.