Distributional Learning of Variational AutoEncoder: Application to Synthetic Data Generation

The Gaussianity assumption has been consistently criticized as a main limitation of the Variational Autoencoder (VAE) despite its efficiency in computational modeling. In this paper, we propose a new approach that expands the model capacity (i.e., expressive power of distributional family) without sacrificing the computational advantages of the VAE framework. Our VAE model's decoder is composed of an infinite mixture of asymmetric Laplace distribution, which possesses general distribution fitting capabilities for continuous variables. Our model is represented by a special form of a nonparametric M-estimator for estimating general quantile functions, and we theoretically establish the relevance between the proposed model and quantile estimation. We apply the proposed model to synthetic data generation, and particularly, our model demonstrates superiority in easily adjusting the level of data privacy.

Paper

Similar papers

Peer review

Reviewer 4gp66/10 · confidence 3/52023-07-03

Summary

In this paper, the authors present a novel nonparametric distributional learning method for VAE. They provide a comprehensive theoretical analysis of the proposed solution (DistVAE), as well as prove its superiority over CTGAN, TVAE, and CTAB-GAN when applied to synthetic data generation on real tabular datasets.

Strengths

1. The theory of the paper seems to be original and significant. 2. The experimental results obtained for DistVAE are superior to those of competitors. 3. The paper is well written.

Weaknesses

1. The experimental setting is limited to tabular datasets and few synthesizers. 2. A more detailed discussion of limitations would be fruitful.

Questions

1. The authors claim that their approach "extends model capacity [...] without sacrificing the computational advantages of the VAE framework". Perhaps the paper would benefit from an experimental comparison (in terms of capturing data distribution and computational efficiency) of the proposed model with different VAEs trained on classical image datasets. Why was this not supplied? My final opinion will depend on the authors' response to this question. 2. Minor comments: l. 14: I believe the authors would like to (additionally) cite another original paper on VAE: Rezende et al., ICML 2014. l. 104: '.' --> ':'. l. 133: $q(x|z; \phi)dx$ --> $q(x|z; \phi)$. l. 136: $q_j$ was already used to denote the number of levels for $x_j$. l. 156: "hessian" --> "Hessian". l. 160: "bounded above" --> "bounded from above".

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

2 fair

Presentation

3 good

Contribution

2 fair

Limitations

A more detailed discussion of limitations would be fruitful.

Reviewer B22e6/10 · confidence 2/52023-07-04

Summary

This paper introduces a novel distributional learning method of VAE, which aims to effectively capture the underlying distribution of the observed dataset using a nonparametric approach. The proposed decoder is composed of an infinite mixture of asymmetric Laplacian distribution, which possesses general distribution fitting capabilities for continuous variables. The authors provide empirical results and demonstrate superiority in easily adjusting the level of data privacy.

Strengths

This work provides concrete theoretical analysis and empirical results with the proposed distributional decoder. The proposed decoder has more capacity to capture the data distribution than the standard Gaussian.

Weaknesses

As [15] indicated, “the model training becomes unstable if the estimated variance of the decoder shrinks to zero in Gaussian VAE” is not theoretically correct. I don't find major weakness in this work.

Questions

Why $\alpha$ is assigned with a prior? I don’t get any intuition from your analysis.

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

2 fair

Contribution

2 fair

Limitations

Too many assumptions in the theoretical analysis which make the generality to real data suspicious.

Reviewer a4Hg6/10 · confidence 2/52023-07-06

Summary

In this paper the authors propose to define the decoder of a VAE using an infinite mixture of asymmetric Laplacian distribution, which results in distributional learning. Specifically, they design a non-parametric generative model in the VAEs. This class of VAEs estimates the conditional CDF while trained with ELBO objective as standard VAEs. They evaluated their approach on multiple datasets and show that their method outperforms baselines.

Strengths

This idea of increasing the model capacity of VAEs is well motivated, because Gaussian assumption has limited the applications of VAEs in the past. So I think this paper brings up some interesting topic and is novel as well. In general this paper is clearly structured. The authors provide detailed derivations of the methodology. Also they perform comprehensive evaluations on multiple datasets.

Weaknesses

My main confusion is about the 2 dimensional latent space in the evaluations. If I understand correctly, all the models are defined with 2D latent space. I wonder if this puts limits on the decoder capacity by itself, and I wonder how the model performs when we do higher latent dimensions in the experiments.

Questions

Please see my questions in the previous section

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

I don't think there is any potential negative impact in their work.

Reviewer 4gp62023-08-16

Thank you for the response

I am satisfied with the authors' rebuttal. I have raised my rating accordingly.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC