Reducing the dimensionality of data using tempered distributions

We reformulate unsupervised dimension reduction problem (UDR) in the language of tempered distributions, i.e. as a problem of approximating an empirical probability density function pemp(x) by another tempered distribution q(x) whose support is in a k-dimensional subspace. Thus, our problem is reduced to the minimization of the distance between q and pemp, D(q,pemp), over a pertinent set of generalized functions. This infinite-dimensional formulation allows to establish a connection with another classical problem of data science --- the sufficient dimension reduction problem (SDR). Thus, an algorithm for the first problem induces an algorithm for the second and vice versa. In order to reduce an optimization problem over distributions to an optimization problem over ordinary functions we introduce a nonnegative penalty function R(f) that ``forces'' the support of f to be k-dimensional. Then we present an algorithm for minimization of I(f)+λR(f), based on the idea of two-step iterative computation, briefly described as a) an adaptation to real data and to fake data sampled around a k-dimensional subspace found at a previous iteration, b) calculation of a new k-dimensional subspace. We demonstrate the method on 4 examples (3 UDR and 1 SDR) using synthetic data and standard datasets.

Paper

Similar papers

© 2026 NYSGPT2525 LLC