K-Means Clustering Optimization Using the Elbow Method and Early Centroid Determination Based on Mean and Median Formula
: The most widely used algorithm in the cluster partitioning method is the K-Means algorithm. Historically K-Means is still the best grouping algorithm among other grouping algorithms with the ability to group a number of data with relatively fast and efficient computing time. The KMeans algorithm is widely implemented in various fields in industrial and scientific applications and is very suitable for processing quantitative data with numeric attributes but there are still weaknesses in this algorithm. Weaknesses of the K-Means algorithm include determining the number of clusters based on assumptions and relying heavily on initial selection of centroids to overcome this weakness, in this study, we propose the use of the elbow method to determine the best number of clusters and determination of centroid based-on mean and median data. The results of this study indicate that using initial centroid determination based on mean data makes the number of iterations needed to achieve uniformity in clusters 22.58% less than using initial random cluster determination and determining the best number of clusters using the elbow method makes the required iteration 25% less than using the number of other clusters.
Paper
Full text
K-Means Clustering Optimization Using the Elbow Method and Early Centroid Determination Based on Mean and Median Formula
Semantic Scholar · Computer Science · 2020
Abstract
: The most widely used algorithm in the cluster partitioning method is the K-Means algorithm. Historically K-Means is still the best grouping algorithm among other grouping algorithms with the ability to group a number of data with relatively fast and efficient computing time. The KMeans algorithm is widely implemented in various fields in industrial and scientific applications and is very suitable for processing quantitative data with numeric attributes but there are still weaknesses in this algorithm. Weaknesses of the K-Means algorithm include determining the number of clusters based on assumptions and relying heavily on initial selection of centroids to overcome this weakness, in this study, we propose the use of the elbow method to determine the best number of clusters and determination of centroid based-on mean and median data. The results of this study indicate that using initial centroid determination based on mean data makes the number of iterations needed to achieve uniformity in clusters 22.58% less than using initial random cluster determination and determining the best number of clusters using the elbow method makes the required iteration 25% less than using the number of other clusters.
References (18)
Scroll for more · 6 remaining