A Comprehensive Survey on Advanced Ancient Document OCR

This paper reviews recent advancements in Optical Character Recognition (OCR) for ancient documents to establish a robust digital archiving framework. Historical records present challenges including physical degradation, non-standard layouts with complex interlinear notes, and a scarcity of labeled data. We delineate the paradigm shift from local feature extraction in CNNs to global context modeling using Transformers and Vision-Language models. For preprocessing, structural restoration via Hierarchical Networks and quality enhancement through Diffusion-based binarization are examined. Layout analysis techniques, such as DBNet for non-linear text flow and Graph Convolutional Networks for logical structure recovery, are evaluated to ensure hierarchical integrity. To overcome data scarcity, we explore diverse learning paradigms: weakly-, semi-, and self-supervised learning (e.g., Masked Autoencoders), along with transfer learning for domain adaptation. Furthermore, we review specialized datasets like HUST-OBC and AI Hub that facilitate effective model benchmarking. This review serves as a technical roadmap for improving the robustness of East Asian document digitization, bridging the gap between historical preservation and advanced machine intelligence.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC