IDENTIFYING MULTIPLE LANGUAGES IN A CONTENT ITEM

Patent №

US 10,180,935

Granted

2019-01-15

Filed 2017

Owner

FACEBOOK, INC.

AI components

5

ml · nlp · speech · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

15422463

A system for identifying language(s) for content items is disclosed. The system can identify different languages for content item words segments by identifying segment languages that maximize a probability across the segments. The probability can be a combination of: an author's likelihood for the language identified for the first word; a combination of transition frequencies for selected languages identified for words, the transition frequencies indicating likelihoods that a transition occurred to the selected language from the previous word's language; and a combination of observation probabilities indicating, for a given word in the content item, a likelihood the given word is in the identified language. For an in-vocabulary word, the observation probabilities can be based on learned probability for that word. For an out-of-vocabulary word, the probability can be computed by breaking the word into overlapping n-grams and computing combined learned probabilities that each n-gram is in the given language.

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
AI hardware1.00
Knowledge representation0.84
Vision0.17
Evolutionary computation0.08
Planning0.04

Ownership

FACEBOOK, INC.

assignment · 415680069

Assignors

FUNIAK, STANISLAV, MERL, DANIEL MATTHEW, PAL, ADITYA, PARK, SEYOUNG, HUANG, FEI, HERDAGDELEN, AMAC

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC