Patent №
US 12,694,702
Granted
2026-07-28
Filed 2024
Owner
Intuit Inc.
Lab
—
AI components
0
Assignment
None on record
Dataset
AIPD
Application
18410785
Systems and methods are disclosed for training an electronic document extraction model, including the generation of the training data to train the model based on sampling a pool of electronic documents based on a rareness metric of the documents. Each electronic document has a document-level rareness metric generated, with the document-level rareness metric being based on one or more of a structural rareness metric or a content rareness metric of the document. The structural rareness metric measures the rareness of the document structure, which may be irrespective of the text content of the document. The content rareness metric measures the rareness of the document content, which may be irrespective of the document structure. The electronic documents are sampled based on the document-level rareness metrics to increase the number of rare documents in the training data without unduly biasing the sampling to optimize the training data for training the extraction model.
Ownership
Intuit Inc.