Patent №
US 9,348,809
Granted
2016-05-24
Filed 2015
Owner
LINKEDIN CORPORATION
Lab
—
AI components
6
ml · nlp · speech · kr · planning · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
14611816
Techniques for training a tokenizer (or word segmenter) are provided. In one technique, a tokenizer tokenizes a token string to identify individual tokens or words. A language model is generated based on the identified tokens or words. A vocabulary about an entity, such as a person or company, is identified. The vocabulary may be online data that refers to the entity, such as a news article or a profile page of a member of a social network. Some of the tokens in the vocabulary may be weighted higher than others. The language model accepts the weighted vocabulary as input and generates pseudo sentences. Alternatively, regular expressions are used to generate the pseudo sentences. The pseudo sentences are used to train the tokenizer.
AI classification
Ownership
LINKEDIN CORPORATION
assignment · 348680291
Assignors
ZHAO, BING, ZHANG, ETHAN
On an employer assignment, the assignors are typically the inventors.