In this paper, we assume the decomposition of points with the mixture of Gaussian distributions in each dimension as an underlying assumption for feature formation of input data. The new guideline presents a unified approach to current basic assumptions and also provides us with an opportunity to solve an essential problem of low-level clustering algorithms. The issue is in the form of the curse of dimensionality which claims that multivariate clustering is meaningless for high dimensional data. To solve this problem, we propose a new type of vector norm (||._||_c) and subsequently Clustering Distance (CD) which is a distance metric system that guarantees meaningfulness even in high dimensional data. The experiments on synthetic and non-synthetic datasets show the effectiveness of the proposed method compared to the current solutions.
Paper
Full text
Meaningful Distance for Multivariate Clustering
Semantic Scholar · Computer Science · 2018
Abstract
In this paper, we assume the decomposition of points with the mixture of Gaussian distributions in each dimension as an underlying assumption for feature formation of input data. The new guideline presents a unified approach to current basic assumptions and also provides us with an opportunity to solve an essential problem of low-level clustering algorithms. The issue is in the form of the curse of dimensionality which claims that multivariate clustering is meaningless for high dimensional data. To solve this problem, we propose a new type of vector norm (||._||_c) and subsequently Clustering Distance (CD) which is a distance metric system that guarantees meaningfulness even in high dimensional data. The experiments on synthetic and non-synthetic datasets show the effectiveness of the proposed method compared to the current solutions.