The challenge of enterprise document processing is in the ability to work with various formats, layouts, and variations of quality of business documents. A multimodal architecture described in this paper is production ready and enables the use of transformer-based, optical character recognition alongside the Gemini API to extract automated invoice data. The hybrid processing pipeline proposed is a combination of the flexibility of the visual-textual reasoning and the deterministic validation of the schema, to increase the accuracy and reliability. Our system is tested on a large collection of 300 business invoices containing digital PDFs and scanned documents. Findings show that the extraction accuracy is 0.95 at the field-level with F1 values of 0.95 and 0.87 on digital invoices and scanned documents respectively, and that the schema is strictly followed. It incorporates confidence-based routing into the architecture and 92 per cent automation rate on high-quality documents. Our solution serves the major needs of an enterprise such as auditability, cost optimization, and scalable deployment, with the solution being an important upgrade over the conventional template-based extraction systems.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex