Patent №
US 8,463,598
Granted
2013-06-11
Filed 2011
Owner
—
Lab
—
AI components
5
ml · nlp · speech · kr · hardware
Assignment
None on record
Dataset
AIPD
2023_r1 edition
Application
13016338
Methods, systems, and apparatus, including computer program products, in which data from web documents are partitioned into a training corpus and a development corpus are provided. First word probabilities for words are determined for the training corpus, and second word probabilities for the words are determined for the development corpus. Uncertainty values based on the word probabilities for the training corpus and the development corpus are compared, and new words are identified based on the comparison.