EFFICIENT LANGUAGE IDENTIFICATION

Patent №

US 8,027,832

Granted

2011-09-27

Filed 2005

Owner

MICROSOFT CORPORATION

AI components

5

ml · nlp · vision · speech · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11056707

A system and methods of language identification of natural language text are presented. The system includes stored expected character counts and variances for a list of characters found in a natural language. Expected character counts and variances are stored for multiple languages to be considered during language identification. At run-time, one or more languages are identified for a text sample based on comparing actual and expected character counts. The present methods can be combined with upstream analyzing of Unicode ranges for characters in the text sample to limit the number of languages considered. Further, n-gram methods can be used in downstream processing to select the most probable language from among the languages identified by the present system and methods.

Machine learningNatural languageVisionSpeechAI hardwareG06F 40/263Y10T 70/30Y10T 70/358Y10T 70/5726Y10T 70/5973Y10T 70/7057

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
Vision0.89
AI hardware0.75
Knowledge representation0.29
Evolutionary computation0.00
Planning0.00

Ownership

MICROSOFT CORPORATION

assignment · 158090440

Assignors

RAMSEY, WILLIAM D., SCHMID, PATRICIA M., POWELL, KEVIN R.

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC