ELECTRONIC DOCUMENT SOURCE INGESTION FOR NATURAL LANGUAGE PROCESSING SYSTEMS

Patent №

US 9,053,085

Granted

2015-06-09

Filed 2012

Owner

INTERNATIONAL BUSINESS MACHINES CORPORATION

AI components

4

nlp · speech · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

13709413

The data store for a natural-language computing system may include information that originates from a plurality of different data sources—e.g., journals, websites, magazines, reference books, and the like. In one embodiment, the information or text from the data sources are converted into a single, shared format and stored as objects in a data store. In order to ingest the different documents with their respective formats, a natural language processing system may perform preprocessing to change the different formats into a normalized format. When a new text document is received, the text may be correlated to a particular properties file which includes instructions specifying how the preprocessor should interpret the received text. Based on these instructions, a preprocessor identifies relevant portions of the text document and assigns these portions to formatting elements in the normalized format. The text may then be stored in the objects based on this assignment.

AI classification

Natural language1.00
Knowledge representation1.00
Speech1.00
AI hardware0.97
Vision0.41
Machine learning0.24
Evolutionary computation0.00
Planning0.00

Ownership

INTERNATIONAL BUSINESS MACHINES CORPORATION

assignment · 294360452

Assignors

DUBBELS, JOEL C.

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC