Clustering in High-Dimension – Tools and Challenges

Dimensionality reduction methods such as Multidimensional scaling (MDS) or t-Distributed Stochastic Neighbor Embedding ( t-SNE) are often followed by clustering in the reduced plot. To examine whether or to what extent these methods affect clustering, we simulate several data structures and apply clustering methods. We first perform clustering using the data in the original space, where we know the true clusters, then perform MDS and t-SNE to scale the data down to two dimensions, cluster on this projected data, and compare differences in the results. We find that MDS and t-SNE can either increase or decrease clustering performance, and are unable to correctly represent data structures with certain shape structures, or in the presence of noise. We examine several clustering methods and show that their performance depends to a large extent on the structure of the data, original dimension, and the noise level, even before we perform dimensionality reduction via MDS and t-SNE. No method among the ones considered here dominates the others in terms of clustering accuracy. KEYWORDS: K-means; Hierarchical clustering; Ward linkage; Complete linkage; Average linkage; Single linkage; Persistence Diagram; Simulation

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC