When clustering data, the importance of attributes affects the final result. The traditional system clustering method only performs clustering based on Euclidean distance, and simply calculates the geometric similarity between data, while ignoring different attributes. In this paper, the generalized Euclidean distance method is used to modify the system clustering algorithm. The rough set theory is applied to the determination of attribute weights, and a complete set of attribute weight calculation process is constructed. The algorithm is easy to implement and can eliminate redundant information. For the first time, mutual information entropy was used to evaluate the performance of the clustering algorithm, and the performance analysis was carried out together with the Friedman test. The data in the common data packet UCI is selected for simulation verification. The conclusion that the performance of the algorithm is significantly improved can be obtained from both qualitative and quantitative aspects.
Paper
Full text
Improved System Cluster Analysis Based on Rough Set and General Euclidean Distance
Semantic Scholar · Computer Science · 2019
Abstract
When clustering data, the importance of attributes affects the final result. The traditional system clustering method only performs clustering based on Euclidean distance, and simply calculates the geometric similarity between data, while ignoring different attributes. In this paper, the generalized Euclidean distance method is used to modify the system clustering algorithm. The rough set theory is applied to the determination of attribute weights, and a complete set of attribute weight calculation process is constructed. The algorithm is easy to implement and can eliminate redundant information. For the first time, mutual information entropy was used to evaluate the performance of the clustering algorithm, and the performance analysis was carried out together with the Friedman test. The data in the common data packet UCI is selected for simulation verification. The conclusion that the performance of the algorithm is significantly improved can be obtained from both qualitative and quantitative aspects.