DETECTION AND EXTRACTION OF ELEMENTS CONSTITUTING IMAGES IN UNSTRUCTURED DOCUMENT FILES

Patent №

US 8,645,819

Granted

2014-02-04

Filed 2011

Owner

XEROX CORPORATION

Lab

AI components

3

nlp · vision · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

13162858

A method and a system for detecting and extracting images in an electronic document are disclosed. The method includes receiving an electronic document and identifying elements of a page. The identified elements include a set of graphical elements and a set of text elements. The method may include identifying and excluding elements which serve as graphical page constructs and/or text formatting elements. The page can then be segmented, based on (remaining) graphical elements and identified white spaces, to generate a set of image blocks. Text elements that are associated with a respective image block are identified as captions. Overlapping candidate images are then grouped to form a new image. The new image can thus include candidate images which would, without the identification of their caption(s), each be treated as a respective image.

AI classification

Vision1.00
AI hardware0.99
Natural language0.72
Evolutionary computation0.01
Machine learning0.00
Planning0.00
Knowledge representation0.00
Speech0.00

Ownership

XEROX CORPORATION

assignment · 264580951

Assignors

DEJEAN, HERVE

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC