Optical Character Recognition of 19th Century Classical Commentaries: the Current State of Affairs
Together with critical editions and translations, commentaries are one of the\nmain genres of publication in literary and textual scholarship, and have a\ncentury-long tradition. Yet, the exploitation of thousands of digitized\nhistorical commentaries was hitherto hindered by the poor quality of Optical\nCharacter Recognition (OCR), especially on commentaries to Greek texts. In this\npaper, we evaluate the performances of two pipelines suitable for the OCR of\nhistorical classical commentaries. Our results show that Kraken + Ciaconna\nreaches a substantially lower character error rate (CER) than Tesseract/OCR-D\non commentary sections with high density of polytonic Greek text (average CER\n7% vs. 13%), while Tesseract/OCR-D is slightly more accurate than Kraken +\nCiaconna on text sections written predominantly in Latin script (average CER\n8.2% vs. 8.4%). As part of this paper, we also release GT4HistComment, a small\ndataset with OCR ground truth for 19th classical commentaries and Pogretra, a\nlarge collection of training data and pre-trained models for a wide variety of\nancient Greek typefaces.\n
Paper
References (27)
Scroll for more · 15 remaining