Information extraction from unstructured documents, meant only for human readers, has to be dealt with differently than from the structured documents. Unstructured documents include visual clues that draw human attention and convey the majority of information to readers. There have been several recent advancements in information extraction in such documents using the conventional natural language processing methodologies. However, there has been little to no work towards using the non-sequential relationships that are found only in unstructured documents for the task of information extraction. In this study, we propose novel methodologies to capture the non-sequential relationships present in the unstructured documents for the task of Named Entity Recognition (NER) using Conditional Random Field (CRF). We experiment with two different datasets having different types of logical reading order and we compare three sets of features. The NER model, that uses the proposed novel features, achieves mean F1-Scores of 68.15% on Retail Receipt and 85.54% on Air Ticket documents.
Paper
Full text
Effectiveness of Visual Features on Diverse Reading Orders for Information Extraction
Semantic Scholar · Computer Science · 2019
Abstract
Information extraction from unstructured documents, meant only for human readers, has to be dealt with differently than from the structured documents. Unstructured documents include visual clues that draw human attention and convey the majority of information to readers. There have been several recent advancements in information extraction in such documents using the conventional natural language processing methodologies. However, there has been little to no work towards using the non-sequential relationships that are found only in unstructured documents for the task of information extraction. In this study, we propose novel methodologies to capture the non-sequential relationships present in the unstructured documents for the task of Named Entity Recognition (NER) using Conditional Random Field (CRF). We experiment with two different datasets having different types of logical reading order and we compare three sets of features. The NER model, that uses the proposed novel features, achieves mean F1-Scores of 68.15% on Retail Receipt and 85.54% on Air Ticket documents.