MODIFYING A TOKENIZER BASED ON PSEUDO DATA FOR NATURAL LANGUAGE PROCESSING

Patent №

US 9,348,809

Granted

2016-05-24

Filed 2015

Owner

LINKEDIN CORPORATION

Lab

AI components

6

ml · nlp · speech · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

14611816

Techniques for training a tokenizer (or word segmenter) are provided. In one technique, a tokenizer tokenizes a token string to identify individual tokens or words. A language model is generated based on the identified tokens or words. A vocabulary about an entity, such as a person or company, is identified. The vocabulary may be online data that refers to the entity, such as a news article or a profile page of a member of a social network. Some of the tokens in the vocabulary may be weighted higher than others. The language model accepts the weighted vocabulary as input and generates pseudo sentences. Alternatively, regular expressions are used to generate the pseudo sentences. The pseudo sentences are used to train the tokenizer.

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
Planning1.00
Knowledge representation1.00
AI hardware0.77
Vision0.05
Evolutionary computation0.00

Ownership

LINKEDIN CORPORATION

assignment · 348680291

Assignors

ZHAO, BING, ZHANG, ETHAN

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC