SYSTEMS AND METHODS FOR AUTOMATIC DATA EXTRACTION FROM DOCUMENT IMAGES

Patent №

US 11,328,524

Granted

2022-05-10

Filed 2019

Owner

UIPATH SRL

Lab

AI components

4

ml · nlp · vision · kr

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16504838

Described systems and methods allow the automatic extraction of structured information from images of structured text documents such as invoices and receipts. Some embodiments employ optical character recognition (OCR) technology to extract individual text tokens (e.g., words) and token bounding boxes from a document image. A feature vector of each text token comprises a first part determined according to a character content of the text token, and a second part determined according to an image content of the token's bounding box. A neural network classifier produces a label indicative of a type of information (e.g. “billing address”, “due date”, etc.) carried by each text token. In some embodiments, documents are linearized by ordering text tokens in a sequence according to a reading order of a natural language (e.g., English, Arabic) in which the respective document is formulated. Token feature vectors are fed to the classifier in the order indicated by the token sequence.

Machine learningNatural languageVisionKnowledge representationG06F 16/56G06F 16/5846G06F 40/284G06V 10/82G06V 30/19173G06V 30/412G06V 30/413G06V 30/414+2 more

AI classification

Vision1.00
Natural language1.00
Machine learning1.00
Knowledge representation0.99
Planning0.44
Speech0.43
AI hardware0.09
Evolutionary computation0.00

Ownership

UIPATH SRL

assignment · 500020528

Assignors

CRISTESCU, HORIA, ADAM, STEFAN A., NEAGOVICI, MIRCEA

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC