CAS Condensed and Accelerated Silhouette: An Efficient Method for Determining the Optimal K in K-Means Clustering

Clustering is a critical component of decision-making in todays data-driven environments. It has been widely used in a variety of fields such as bioinformatics, social network analysis, and image processing. However, clustering accuracy remains a major challenge in large datasets. This paper presents a comprehensive overview of strategies for selecting the optimal value of k in clustering, with a focus on achieving a balance between clustering precision and computational efficiency in complex data environments. In addition, this paper introduces improvements to clustering techniques for text and image data to provide insights into better computational performance and cluster validity. The proposed approach is based on the Condensed Silhouette method, along with statistical methods such as Local Structures, Gap Statistics, Class Consistency Ratio, and a Cluster Overlap Index CCR and COIbased algorithm to calculate the best value of k for K-Means clustering. The results of comparative experiments show that the proposed approach achieves up to 99 percent faster execution times on high-dimensional datasets while retaining both precision and scalability, making it highly suitable for real time clustering needs or scenarios demanding efficient clustering with minimal resource utilization.

Paper

References (23)

04For each batch, the function func is applied independently
05The dataset X is split into batches of size b , determined either statically or dynamically
06“An information-theoretic approach to k selection in k-means clustering,”Information Sciences
07“Methods for optimal k selection in k-means clustering: A comparative review,”Journal of Machine Learning Research
08“Optimal clusters in k-means: Balancing efficiency and accuracy for big data,”Data Mining and Knowledge Discovery
09“Optimal k selection in large-scale k-means clustering,”International Journal of Data Science
10“Determining k with density-based approaches in k-means,”Journal of Artificial Intelligence Research
11“Evaluating k in k-means with elbow and gap statistic methods,”IEEE Transactions on Big Data
12“Determining optimal k in streaming data,”Journal of Real-Time Data Science

Scroll for more · 11 remaining

Similar papers

© 2026 NYSGPT2525 LLC