An Evaluation of DNN Architectures for Page Segmentation of Historical Newspapers

One important and particularly challenging step in the optical character\nrecognition (OCR) of historical documents with complex layouts, such as\nnewspapers, is the separation of text from non-text content (e.g. page borders\nor illustrations). This step is commonly referred to as page segmentation.\nWhile various rule-based algorithms have been proposed, the applicability of\nDeep Neural Networks (DNNs) for this task recently has gained a lot of\nattention. In this paper, we perform a systematic evaluation of 11 different\npublished DNN backbone architectures and 9 different tiling and scaling\nconfigurations for separating text, tables or table column lines. We also show\nthe influence of the number of labels and the number of training pages on the\nsegmentation quality, which we measure using the Matthews Correlation\nCoefficient. Our results show that (depending on the task) Inception-ResNet-v2\nand EfficientNet backbones work best, vertical tiling is generally preferable\nto other tiling approaches, and training data that comprises 30 to 40 pages\nwill be sufficient most of the time.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC