VisualWordGrid: Information Extraction From Scanned Documents Using A Multimodal Approach

We introduce a novel approach for scanned document representation to perform\nfield extraction. It allows the simultaneous encoding of the textual, visual\nand layout information in a 3-axis tensor used as an input to a segmentation\nmodel. We improve the recent Chargrid and Wordgrid \\cite{chargrid} models in\nseveral ways, first by taking into account the visual modality, then by\nboosting its robustness in regards to small datasets while keeping the\ninference time low. Our approach is tested on public and private document-image\ndatasets, showing higher performances compared to the recent state-of-the-art\nmethods.\n

Paper

References (20)

Scroll for more · 8 remaining

Similar papers

© 2026 NYSGPT2525 LLC