METHODS AND APPARATUS FOR RAPID ACOUSTIC UNIT SELECTION FROM A LARGE SPEECH CORPUS

Patent №

US 8,315,872

Granted

2012-11-20

Filed 2011

Owner

AT&T CORP.

Lab

AI components

2

nlp · speech

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

13306157

A speech synthesis system can select recorded speech fragments, or acoustic units, from a very large database of acoustic units to produce artificial speech. The selected acoustic units are chosen to minimize a combination of target and concatenation costs for a given sentence. However, as concatenation costs, which are measures of the mismatch between sequential pairs of acoustic units, are expensive to compute, processing can be greatly reduced by pre-computing and caching the concatenation costs. Unfortunately, the number of possible sequential pairs of acoustic units makes such caching prohibitive. However, statistical experiments reveal that while about 85% of the acoustic units are typically used in common speech, less than 1% of the possible sequential pairs of acoustic units occur in practice. A method for constructing an efficient concatenation cost database is provided by synthesizing a large body of speech, identifying the acoustic unit sequential pairs generated and their respective concatenation costs, and storing those concatenation costs likely to occur. By constructing a concatenation cost database in this faction, the processing power required at run-time is greatly reduced with negligible effect on speech quality.

Natural languageSpeechG10L 13/07G10L 13/00G10L 13/027G10L 13/08

AI classification

Natural language1.00
Speech1.00
Machine learning0.49
Evolutionary computation0.08
AI hardware0.05
Vision0.03
Planning0.02
Knowledge representation0.01

Ownership

AT&T CORP.

assignment · 272920981

Assignors

BEUTNAGEL, MARK CHARLES, MOHRI, MEHRYAR, RILEY, MICHAEL DENNIS

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC