We revisit the classical problem of estimating an unknown distribution from its samples by fitting a mixture model that minimizes cross-entropy loss. Framing the task as a stochastic convex optimization problem over the space of M-component mixture distributions, we propose a family of estimators derived from the stochastic mirror descent (SMD) algorithm. This optimization-based approach provides a principled and flexible framework that generalizes traditional estimators and accommodates a variety of geometries through the choice of Bregman divergences.A key advantage of our method is that it scales efficiently with the number of candidate components fi; that is, one can employ a large set of basis distributions in the mixture model without incurring significant computational overhead. This enables richer approximations and improved estimation accuracy.Moreover, unlike methods that require strict lower bounds on mixture weights (i.e., restricting to the simplex to a denser version where each component is at least δ >0), our framework operates over the full probability simplex ΔM. This allows the estimator to naturally suppress irrelevant components, yielding sparser solutions when appropriate without hard-thresholding.We demonstrate that, under mild conditions, the proposed φ-SMD estimators achieve near-optimal convergence rates in both Kullback–Leibler (KL) divergence and ℓ2-norm, offering practical benefits particularly in high-dimensional or computationally constrained scenarios. Our analysis highlights improved performance guarantees over classical estimators, particularly in terms of sample efficiency, scalability, and dimensional dependence.
Paper
References (22)
Scroll for more · 10 remaining