SYSTEM AND METHOD FOR SELECTING TRAINING TEXT

Patent №

US 6,038,533

Granted

2000-03-14

Filed 1995

Owner

AT&T IPM CORP.

Lab

AI components

5

ml · nlp · vision · speech · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

08499159

A system and method are described for determining a near-optimum subset of data, based on a selected model, from a large corpus of data. Sets of feature vectors corresponding to natural or other preselected divisions of the data corpus are mapped into matrices representative of such divisions. The invention operates to find a submatrix of full rank formed as a union of one or more of those division-based matrices. A greedy algorithm utilizing Gram-Schmidt orthonormalization operates on the division matrices to find a near optimum submatrix and in a time bound representing a substantial improvement over prior-art methods. An important application of the invention is the selection of a small number of sentences from a corpus of a very large number of such sentences from which the parameters of a duration model for speech synthesis can be estimated.

AI classification

Speech1.00
Natural language1.00
Vision0.99
Machine learning0.98
AI hardware0.82
Evolutionary computation0.29
Planning0.05
Knowledge representation0.00

Ownership

AT&T IPM CORP.

assignment · 75830324

Assignors

BUCHSBAUM, ADAM LOUIS, VAN SANTEN, JAN PIERTER

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC