SYSTEM AND METHOD FOR EXTRACTING ENTITIES OF INTEREST FROM TEXT USING N-GRAM MODELS

Patent №

US 7,493,293

Granted

2009-02-17

Filed 2006

Owner

INTERNATIONAL BUSINESS MACHINES CORPORATION

AI components

6

ml · nlp · vision · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11421379

A document (or multiple documents) is analyzed to identify entities of interest within that document. This is accomplished by constructing n-gram or bi-gram models that correspond to different kinds of text entities, such as chemistry-related words and generic English words. The models can be constructed from training text selected to reflect a particular kind of text entity. The document is tokenized, and the tokens are run against the models to determine, for each token, which kind of text entity is most likely to be associated with that token. The entities of interest in the document can then be annotated accordingly.

AI classification

Natural language1.00
Machine learning1.00
Planning0.93
Vision0.93
Knowledge representation0.87
AI hardware0.81
Speech0.48
Evolutionary computation0.00

Ownership

INTERNATIONAL BUSINESS MACHINES CORPORATION

assignment · 181580559

Assignors

KANUNGO, TAPAS, RHODES, JAMES J

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC