During the last few years, research in historical document restoration and understanding (HDRU) has gained increasing popularity. One major problem facing HDRU is the presence of degradation, which renders historical documents unreadable. Although promising results have been obtained, these methods lack the ability to generalize across different datasets. Also, multiple pre-processing and post-processing steps are used, which add more computational complexity and make inference unpractical in real-life settings. Deep Learning (DL) has been successfully used to solve various supervised learning problems in computer vision, where labeled datasets are readily available. However, in HDRU, large annotated historical document datasets are not available. In this paper, we propose an efficient multitask learning (MTL) approach that is based on jointly training self-supervised and supervised learning modules. In the self-supervised learning module, we define two tasks that can be trained with unlabeled data. The first task consists of denoising, and the second task is to learn handwritten characteristics (text orientation). In the supervised learning module, a small subset of labeled data is used to perform text extraction or binarization. All the tasks are formulated as a multi-objective Frank-Wolfe-based optimization problem. We show that convergence to a Pareto optimal solution of jointly training multiple tasks together improves the overall invariance and accuracy of the model. DIBCO 2010-2017 datasets were used for training and DIBCO 2018 for testing. We achieved state-of-the-art results with an F-Score measure of 91.21.
Paper
Full text
Frank-Wolfe-based Multi-task Learning for Historical Document Restoration
OpenAlex · Handwritten Text Recognition Techniques · 2022
Abstract
During the last few years, research in historical document restoration and understanding (HDRU) has gained increasing popularity. One major problem facing HDRU is the presence of degradation, which renders historical documents unreadable. Although promising results have been obtained, these methods lack the ability to generalize across different datasets. Also, multiple pre-processing and post-processing steps are used, which add more computational complexity and make inference unpractical in real-life settings. Deep Learning (DL) has been successfully used to solve various supervised learning problems in computer vision, where labeled datasets are readily available. However, in HDRU, large annotated historical document datasets are not available. In this paper, we propose an efficient multitask learning (MTL) approach that is based on jointly training self-supervised and supervised learning modules. In the self-supervised learning module, we define two tasks that can be trained with unlabeled data. The first task consists of denoising, and the second task is to learn handwritten characteristics (text orientation). In the supervised learning module, a small subset of labeled data is used to perform text extraction or binarization. All the tasks are formulated as a multi-objective Frank-Wolfe-based optimization problem. We show that convergence to a Pareto optimal solution of jointly training multiple tasks together improves the overall invariance and accuracy of the model. DIBCO 2010-2017 datasets were used for training and DIBCO 2018 for testing. We achieved state-of-the-art results with an F-Score measure of 91.21.