This paper investigates an effective method to automatically visualize a given data set based on machine learning. Basically, the visualization results can be varied according to the purpose of the data analysis, and as the understanding of the data becomes larger, more various results can be obtained. This paper aims at realization of an automatic data visualization system based on machine learning, and introduces a meta-level feature engineering process to construct a visualization recommendation (classification) model. Through various experiments, we have designed various meta- feature variables to determine the significance of the visualization results in order to develop the automatic visualization system and constructed the visualization recommendation model using the meta-features. For performance evaluation, we have used three data sources including UCI ML Repository, Data.world, and R datasets, and have found that the decision tree-based recommendation model provides the best performance.
Paper
Full text
Machine Learning-based Automated Data Visualization: A Meta-feature Engineering Approach
Semantic Scholar · Computer Science · 2019
Abstract
This paper investigates an effective method to automatically visualize a given data set based on machine learning. Basically, the visualization results can be varied according to the purpose of the data analysis, and as the understanding of the data becomes larger, more various results can be obtained. This paper aims at realization of an automatic data visualization system based on machine learning, and introduces a meta-level feature engineering process to construct a visualization recommendation (classification) model. Through various experiments, we have designed various meta- feature variables to determine the significance of the visualization results in order to develop the automatic visualization system and constructed the visualization recommendation model using the meta-features. For performance evaluation, we have used three data sources including UCI ML Repository, Data.world, and R datasets, and have found that the decision tree-based recommendation model provides the best performance.