TRAINING OF AN ELECTRONIC DOCUMENT EXTRACTION MODEL

Patent №

US 12,694,702

Granted

2026-07-28

Filed 2024

Owner

Intuit Inc.

Lab

AI components

0

Assignment

None on record

Dataset

AIPD

Application

18410785

Systems and methods are disclosed for training an electronic document extraction model, including the generation of the training data to train the model based on sampling a pool of electronic documents based on a rareness metric of the documents. Each electronic document has a document-level rareness metric generated, with the document-level rareness metric being based on one or more of a structural rareness metric or a content rareness metric of the document. The structural rareness metric measures the rareness of the document structure, which may be irrespective of the text content of the document. The content rareness metric measures the rareness of the document content, which may be irrespective of the document structure. The electronic documents are sampled based on the document-level rareness metrics to increase the number of rare documents in the training data without unduly biasing the sampling to optimize the training data for training the extraction model.

G06V 30/19147G06V 30/414G06V 30/412

Ownership

Intuit Inc.

© 2026 NYSGPT2525 LLC