LANGUAGE AUTODETECTION FROM NON-CHARACTER SUB-TOKEN SIGNALS

Patent №

US 11,630,951

Granted

2023-04-18

Filed 2022

Owner

MICROSOFT TECHNOLOGY LICENSING, LLC

AI components

4

ml · nlp · speech · kr

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

17839330

In non-limiting examples of the present disclosure, systems, methods and devices for determining a language of a text string are presented. A language detection model may be maintained. The language detection model may comprise identities and weights for initial and final consonants, identities and weights for prefixes and suffixes, and identities and weights for vowel sequences, where each identity is derived from a training corpus. The weights may correspond to a frequency of a text unit in the corpus. A text string may be received and a match score between the text string and the language of the language detection model may be determined. The match score may be based on initial and final consonant scores, prefix and suffix scores, and/or vowel sequence scores for each word in the text string. If the match score meets a threshold value a follow-up action associated with the language may be performed.

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
Knowledge representation0.76
Vision0.09
AI hardware0.04
Evolutionary computation0.00
Planning0.00

Ownership

MICROSOFT TECHNOLOGY LICENSING, LLC

assignment · 601870017

Assignors

GLASS, ANDREW STUART, MAGNUS, MARGARET HOPE, RADTKE, ROLAND

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC