Software Defect Prediction on Unlabelled Dataset with Machine Learning Techniques

In this work, we are going to explore how to conduct defect software prediction on unlabelled datasets by exploiting unsupervised machine learning techniques. In previous literature, various approaches have been proposed over time with the aim of labelling an unlabelled dataset (i.e. classifying dataset instances/modules in terms of their defectiveness), they can usually rely on other software datasets, software experts or metrics’ thresholds. In this study, we intend to show the results obtained by exploiting CLAMI and CLAMI+, two approaches that overcome the various limitation of the previous ones adopted by researches, since they are independent on metrics’ threshold, do not need experts software knowledge and can be easily automated.Our prediction model takes as input a set of unlabelled datasets and machine learning techniques. Its output is composed of the values of several performance indicators on training datasets and predictions on test datasets. The latter can be employed to deduce information on the status of software code and, consequently, concentrate the software developers’ effort only where necessary.

Paper

Full text

PDF

Software Defect Prediction on Unlabelled Dataset with Machine Learning Techniques

Semantic Scholar · Computer Science · 2019

Abstract

In this work, we are going to explore how to conduct defect software prediction on unlabelled datasets by exploiting unsupervised machine learning techniques. In previous literature, various approaches have been proposed over time with the aim of labelling an unlabelled dataset (i.e. classifying dataset instances/modules in terms of their defectiveness), they can usually rely on other software datasets, software experts or metrics’ thresholds. In this study, we intend to show the results obtained by exploiting CLAMI and CLAMI+, two approaches that overcome the various limitation of the previous ones adopted by researches, since they are independent on metrics’ threshold, do not need experts software knowledge and can be easily automated.Our prediction model takes as input a set of unlabelled datasets and machine learning techniques. Its output is composed of the values of several performance indicators on training datasets and predictions on test datasets. The latter can be employed to deduce information on the status of software code and, consequently, concentrate the software developers’ effort only where necessary.

Similar papers

© 2026 NYSGPT2525 LLC