Semi-supervised K-means++

ABSTRACT Traditionally, practitioners initialize the k-means algorithm with centres chosen uniformly at random. Randomized initialization with uneven weights (k-means++) has recently been used to improve the performance over this strategy in cost and run-time. We consider the k-means problem with semi-supervised information, where some of the data are pre-labelled, and we seek to label the rest according to the minimum cost solution. By extending the k-means++ algorithm and analysis to account for the labels, we derive an improved theoretical bound on expected cost and observe improved performance in simulated and real-data examples. This analysis provides theoretical justification for a roughly linear semi-supervised clustering algorithm.

Paper

Similar papers

© 2026 NYSGPT2525 LLC