The widespread availability of tabular datasets on the internet has facilitated the adoption of technologies for information retrieval and prediction of missing or obscured data. However, a significant proportion of these datasets suffer from poor quality issues, such as misspelled or absent values, incomplete metadata. Furthermore, the manual collection and integration of data from various sources heavily rely on expert users’ knowledge, making the process laborious and time-consuming. Hence, there is a pressing need to streamline the collection and linking of data from multiple sources, ensuring that datasets are readily prepared for the intended analysis. This article introduces an approach that addresses these challenges by offering several key functionalities. Firstly, it facilitates data annotation to address concerns related to misspelled and incomplete metadata. Secondly, it enables data repair to handle missing values within the dataset. Lastly, it provides data augmentation capabilities, allowing the dynamic addition of meaningful columns and their corresponding cell values. The effectiveness of this approach has been evaluated using benchmark datasets with promising results in terms of evaluation metrics.
Paper
Full text
A Journey to Enhance Tabular Data FAIRness : From Annotation to Repair and Augmentation
Semantic Scholar · Computer Science · 2023
Abstract
The widespread availability of tabular datasets on the internet has facilitated the adoption of technologies for information retrieval and prediction of missing or obscured data. However, a significant proportion of these datasets suffer from poor quality issues, such as misspelled or absent values, incomplete metadata. Furthermore, the manual collection and integration of data from various sources heavily rely on expert users’ knowledge, making the process laborious and time-consuming. Hence, there is a pressing need to streamline the collection and linking of data from multiple sources, ensuring that datasets are readily prepared for the intended analysis. This article introduces an approach that addresses these challenges by offering several key functionalities. Firstly, it facilitates data annotation to address concerns related to misspelled and incomplete metadata. Secondly, it enables data repair to handle missing values within the dataset. Lastly, it provides data augmentation capabilities, allowing the dynamic addition of meaningful columns and their corresponding cell values. The effectiveness of this approach has been evaluated using benchmark datasets with promising results in terms of evaluation metrics.