APPARATUS AND METHOD FOR FORMING A FILTERED INFLECTED LANGUAGE MODEL FOR AUTOMATIC SPEECH RECOGNITION
Patent №
US 6,073,091
Granted
2000-06-06
Filed 1997
Owner
IBM CORPORATION
Lab
AI components
4
ml · nlp · speech · kr
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
08906812
A method of forming a language model for a language having a selected vocabulary of word forms comprises: (a) mapping the word forms into integer vectors in accordance with frequencies of word form occurrence; (b) partitioning the integer vectors into subsets, the subsets respectively having ranges of frequencies of word form occurrence associated therewith, the subsets being arranged in a descending order of frequency ranges; (c) respectively assigning maps to the subsets; (d) filtering a textual corpora using the maps assigned to the subsets in order to generate indexed integers; (e) determining n-gram statistics for the indexed integers; and (f) estimating n-gram language model probabilities from the n-gram statistics to form the language model.
AI classification
Ownership
IBM CORPORATION
assignment · 87370627
Assignors
KANEVSKY, DIMITRI, MONKOWSKI, MICHAEL D., SEDIVY, JAN
On an employer assignment, the assignors are typically the inventors.