Abstract In the recent years, the emergence of data mining techniques in the research industries makes it possible to extract valuable information from huge volume of data. Good understanding of data mining techniques is necessary for research scholars and information industries to make use of this opportunity efficiently to improve the quality of their findings. Two most popular techniques of data mining are: clustering and classification. Clustering is an unsupervised learning method that groups the data based on distance and similarities among them. Classification assigns data items into target categories or classes with the aim of predicting the target class for each instance of data in a heterogeneous dataset accurately. This review focuses on current data mining techniques used in classification and clustering of gene expression dataset. Main attention is given to supervise and unsupervised methods that are being used recently to classify gene expression for the purpose of diagnosis and prognosis of terrifying diseases. In fact, gene expression data clustering offers a powerful approach to detect cancers from a given dataset. The various efficient algorithms available in the literature are analyzed to know their availability in the different situations.
Paper
Full text
Mining gene expression data using data mining techniques: A critical review
Semantic Scholar · Computer Science · 2019
Abstract
Abstract In the recent years, the emergence of data mining techniques in the research industries makes it possible to extract valuable information from huge volume of data. Good understanding of data mining techniques is necessary for research scholars and information industries to make use of this opportunity efficiently to improve the quality of their findings. Two most popular techniques of data mining are: clustering and classification. Clustering is an unsupervised learning method that groups the data based on distance and similarities among them. Classification assigns data items into target categories or classes with the aim of predicting the target class for each instance of data in a heterogeneous dataset accurately. This review focuses on current data mining techniques used in classification and clustering of gene expression dataset. Main attention is given to supervise and unsupervised methods that are being used recently to classify gene expression for the purpose of diagnosis and prognosis of terrifying diseases. In fact, gene expression data clustering offers a powerful approach to detect cancers from a given dataset. The various efficient algorithms available in the literature are analyzed to know their availability in the different situations.