TEXT LANGUAGE IDENTIFICATION

Patent №

US 7,689,409

Granted

2010-03-30

Filed 2003

Owner

FRANCE TELECOM

Lab

AI components

2

ml · nlp

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

10732809

After prestoring first character strings that occur frequently in words of languages and second character strings that are a typical therein, a device for automatically identifying the language of a text from a plurality of languages extracts words from the text and constructs all of the character strings contained in each extracted word. Each string in an extracted word is compared to the first and second strings of a particular language. If the word contains a first string, a score of the language is increased by a coefficient depending in particular on the position of the first string in the word. If the word contains a second string, the score is decreased by a coefficient associated with the second string. The highest of the scores corresponding to the predetermined languages identifies the language of the text.

Machine learningNatural languageG06F 40/216G06F 40/263G06F 40/268

AI classification

Natural language1.00
Machine learning0.57
AI hardware0.35
Knowledge representation0.03
Speech0.02
Vision0.00
Evolutionary computation0.00
Planning0.00

Ownership

FRANCE TELECOM

assignment · 151320866

Assignors

HEINECKE, JOHANNES

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC