TableNet: Deep Learning model for end-to-end Table detection and Tabular data extraction from Scanned Document Images

With the widespread use of mobile phones and scanners to photograph and\nupload documents, the need for extracting the information trapped in\nunstructured document images such as retail receipts, insurance claim forms and\nfinancial invoices is becoming more acute. A major hurdle to this objective is\nthat these images often contain information in the form of tables and\nextracting data from tabular sub-images presents a unique set of challenges.\nThis includes accurate detection of the tabular region within an image, and\nsubsequently detecting and extracting information from the rows and columns of\nthe detected table. While some progress has been made in table detection,\nextracting the table contents is still a challenge since this involves more\nfine grained table structure(rows & columns) recognition. Prior approaches have\nattempted to solve the table detection and structure recognition problems\nindependently using two separate models. In this paper, we propose TableNet: a\nnovel end-to-end deep learning model for both table detection and structure\nrecognition. The model exploits the interdependence between the twin tasks of\ntable detection and table structure recognition to segment out the table and\ncolumn regions. This is followed by semantic rule-based row extraction from the\nidentified tabular sub-regions. The proposed model and extraction approach was\nevaluated on the publicly available ICDAR 2013 and Marmot Table datasets\nobtaining state of the art results. Additionally, we demonstrate that feeding\nadditional semantic features further improves model performance and that the\nmodel exhibits transfer learning across datasets. Another contribution of this\npaper is to provide additional table structure annotations for the Marmot data,\nwhich currently only has annotations for table detection.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC