A COMPREHENSIVE SURVEY ON DOCUMENT CLUSTERING TECHNIQUES

Document clustering plays a crucial role in text mining and natural language processing by organizing large volumes of unstructured textual data into meaningful groups without requiring labeled instances. With the rapid growth of digital information, efficient clustering techniques have become essential for tasks such as information retrieval, topic discovery, and knowledge organization. This survey provides a structured analysis of document clustering approaches by categorizing them into three major groups: classical methods, probabilistic models, and nature-inspired optimization techniques. Classical approaches, including K-Means, hierarchical clustering, DBSCAN, and graph-based methods, offer computational efficiency but often struggle to capture deeper semantic relationships. In contrast, probabilistic models such as Latent Dirichlet Allocation (LDA), Gaussian Mixture Models (GMM), and Hidden Markov Models (HMM) aim to uncover latent structures within text data, albeit with increased computational complexity. Furthermore, nature-inspired algorithms-including Genetic Algorithm (GA), Particle Swarm Optimization (PSO), Ant Colony Optimization (ACO), Artificial Bee Colony (ABC), and Grey Wolf Optimization (GWO)-enhance clustering performance through global optimization strategies. The study critically examines the strengths and limitations of these techniques and emphasizes the importance of hybrid and adaptive frameworks to improve clustering accuracy, scalability, and semantic representation.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC