AUTOMATIC EXTRACTION OF HUMAN-READABLE LISTS FROM STRUCTURED DOCUMENTS

Patent №

US 7,558,792

Granted

2009-07-07

Filed 2004

Owner

XEROX CORPORATION

Lab

AI components

5

nlp · vision · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

10879843

One aspect of the invention extracts a human readable list from a document. It does this by accessing a file that contains data that represents a portion of the document. The data is formatted in accordance with a document formatting description. The data is parsed into tokens that include container tokens and textual tokens. From the container tokens, this aspect determines a context for some of the textual tokens. Once the context is determined, this aspect determines a separator pattern between one of the textual tokens and an adjacent textual token where both the textual token and the adjacent textual token have the same context. Once the separator pattern is determined, the textual tokens can be extracted responsive to the separator pattern. Finally, the textual tokens are presented as the human readable list (for example, displayed, returned in a database, returned in response to a function or subroutine call, etc.).

AI classification

Natural language1.00
Planning0.97
Vision0.85
AI hardware0.79
Knowledge representation0.67
Speech0.22
Machine learning0.02
Evolutionary computation0.01

Ownership

XEROX CORPORATION

assignment · 155420952

Assignors

BIER, ERIC A.

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC