Multimodal deep networks for text and image-based document classification

Classification of document images is a critical step for archival of old\nmanuscripts, online subscription and administrative procedures. Computer vision\nand deep learning have been suggested as a first solution to classify documents\nbased on their visual appearance. However, achieving the fine-grained\nclassification that is required in real-world setting cannot be achieved by\nvisual analysis alone. Often, the relevant information is in the actual text\ncontent of the document. We design a multimodal neural network that is able to\nlearn from word embeddings, computed on text extracted by OCR, and from the\nimage. We show that this approach boosts pure image accuracy by 3% on\nTobacco3482 and RVL-CDIP augmented by our new QS-OCR text dataset\n(https://github.com/Quicksign/ocrized-text-dataset), even without clean text\ninformation.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC