On the Consistency of Maximum Likelihood Estimation of Probabilistic Principal Component Analysis

Probabilistic principal component analysis (PPCA) is currently one of the most used statistical tools to reduce the ambient dimension of the data. From multidimensional scaling to the imputation of missing data, PPCA has a broad spectrum of applications ranging from science and engineering to quantitative finance. Despite this wide applicability in various fields, hardly any theoretical guarantees exist to justify the soundness of the maximal likelihood (ML) solution for this model. In fact, it is well known that the maximum likelihood estimation (MLE) can only recover the true model parameters up to a rotation. The main obstruction is posed by the inherent identifiability nature of the PPCA model resulting from the rotational symmetry of the parameterization. To resolve this ambiguity, we propose a novel approach using quotient topological spaces and in particular, we show that the maximum likelihood solution is consistent in an appropriate quotient Euclidean space. Furthermore, our consistency results encompass a more general class of estimators beyond the MLE. Strong consistency of the ML estimate and consequently strong covariance estimation of the PPCA model have also been established under a compactness assumption.

Paper

Similar papers

Peer review

Reviewer QWak3/10 · confidence 3/52023-06-22

Summary

This work studies the consistency of the maximum likelihood estimates of probabilistic PCA. These estimates are unique up to a rotation, so they use the quotient space and claim in their Lemma 6.1, ..., 6.5 stated in Redner[1981] are verified and apply the consistency results available in this book.

Strengths

This paper is overall well written, with a thorough introduction to equivalence classes and quotient topological spaces.

Weaknesses

Compared to Redner [1981] and Wald [1949], the novelty is not enough. I see the contribution as mainly an application of the result of Wald to the case of probabilistic PCA. The general idea is rather straightforward: using the quotient space instead of $R^d$ is the first idea that comes to mind. I agree that the precise formulation is not trivial but alone, and although the writing is excellent, it is not enough for a publication in my opinion.

Questions

Could you clarify the contributions of the paper ? Is there something I missed ?

Rating

3: Reject: For instance, a paper with technical flaws, weak evaluation, inadequate reproducibility and incompletely addressed ethical considerations.

Confidence

3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

1 poor

Limitations

Not applicable.

Reviewer MhxU6/10 · confidence 2/52023-07-04

Summary

In this work, the authors propose a novel topological framework and show that the maximum likelihood (ML) solution of probabilistic principal component analysis (PPCA) is consistent in an appropriate quotient Euclidean space. The consistency results encompass more estimators beyond the ML solution. In addition, the ML solution has been shown to achieve strong consistency when the parameter space is compact.

Strengths

This work seems to be the first work to establish the (strong) consistency result of the ML solution of PPCA.

Weaknesses

It would be ideal to remove the assumption that the latent dimension $q$ is known.

Questions

I am not familiar with the techniques used in this work. Does the consistency in the appropriate quotient Euclidean space means that the ML estimate $(\hat{\mathbf{W}},\hat{\sigma}^2)$ converges to the true parameters $(\mathbf{W}_0, \sigma^2_0)$ in probability up to rotation (i.e., $\hat{\mathbf{W}}$ converges to $\mathbf{W}_0 \mathbf{R}$ with $\mathbf{R}$ being orthogonal)?

Rating

6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations.

Confidence

2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.

Soundness

3 good

Presentation

3 good

Contribution

3 good

Limitations

NA

Reviewer gPBN7/10 · confidence 5/52023-07-06

Summary

The paper discusses consistency of the maximum likelihood (ML) estimation in probabilistic principal component analysis (PPCA). Despite its wide applicability, proving ML estimation consistency in PPCA has been a challenging task because of the non-identifiability of the problem. The author(s) extend the quotient space topology idea presented in Redner (1981) to a more general setting where ML estimator is one such estimator. Also, strong consistency of the ML estimator is proved under a compactness assumption.

Strengths

1. Detailed description of the quotient space topology and the associated metric. The constructive approach was clear and concise. 2. With standard assumptions on the quotient space, the author(s) derived both weak and strong consistency of the ML estimator in the PPCA model. 3. Verification of the Wald's conditions [Wald, 1949] in the quotient space topology with sufficient technical details.

Weaknesses

1. Some remarks on the Wald's conditions on the quotient space to clarify how those conditions could be interpreted in the quotient space would be useful. 2. One or two examples to clarify the result in Theorem 4.2 would be helpful.

Questions

1. What would be $X_{\equiv}$ of Lemma 4.3 in the case of PPCA and in what sense $X_{\equiv}$ and the quotient $X\C$ are "identical" in that lemma? The two metric defined on those two topological spaces are not same- right? 2. I was wondering whether it's not possible to have the consistency result for the MLE in PPCA as a corollary of the results proved in the following Van der Vaart, A. W., & Wellner, J. A. (1992). Existence and consistency of maximum likelihood in upgraded mixture models. Journal of multivariate analysis, 43(1), 133-146. I understand the above may be for mixture models but can't we have a corollary for one mixing component which would be the marginal distribution of the data given W and $\sigma$ in your case.

Rating

7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.

Confidence

5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.

Soundness

3 good

Presentation

4 excellent

Contribution

3 good

Limitations

Yes, the author(s) have adequately addressed the limitations.

Reviewer ZSYo5/10 · confidence 4/52023-07-07

Summary

paper addresses "Probabilistic PCA" which stands for a setup where the vector of p observations x can be written as x = Wz + eps where eps stands for an additive (Gaussian, centered, iid) noise, the matrix W \in R^{p x q} is an unknown and z in R^q is the other unknown . Only the value of q is known in advance and the goal is to recover W and z and remove the noise. The main challenge is to address the fact that the model is invariant w.r.t. applying a rotation R to W and z: Wz = (W R) (R' z) . Note that (W R) & (R' z) have the same euclidean norms as W & z. The contribution is to propose an analysis in the quotient space of the vector space by the rotations.

Strengths

The paper contains a rigorous consistency analysis of the PPCA problem with respect to rotation invariance

Weaknesses

The main focus of the paper on the rotation invariance does not seem to be the biggest concern with PPCA or other related (see Questions) methods.

Questions

1. There is a broader set of problems where rotation invariance causes trouble, and that set of problem directly related to the PPCA problem considered: all problems where the observations / measurements are of the form <u_i, u_j> or <u_i, v_j>. This comprises PPCA, but also matrix factorization and embedding problems. Can authors discuss how their method can expand to those problems? 2. Is there a relation between rotation invariance and computational complexity? Can a rotation be a distractor in the gradient steps convergence of a first order method?

Rating

5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly.

Confidence

4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.

Soundness

3 good

Presentation

4 excellent

Contribution

2 fair

Limitations

The focus seems to be relatively narrow and implications of the proven results not sufficiently emphasized.

Reviewer gPBN2023-08-15

Response to author rebuttal

I thank the author(s) for their detailed explanations in the rebuttal. I am happy with the answers they provided to my queries. I have increased my rating to "accept".

Authorsrebuttal2023-08-16

Correcting a minor typo

Hello everyone, It has come to our notice that there is a minor (and obvious) typo in the response we wrote for reviewer MhxU. In the second line of the rebuttal it should have been $\mathbb{P}(\widehat{W}\widehat{W}^T\to W_0W_0^T)=1$ and not $\mathbb{P}(\widehat{W}\widehat{W}^T\to \widehat{W}_0\widehat{W}_0^T)=1$.

Reviewer QWak2023-08-16

Thanks for your answer. I agree that Redner's work was incomplete and that the current work fills this hole (in my opinion Redner's work is the reason why no one tried to prove consistency of PPCA before). I agree that this method could be used for different models. Nevertheless, once the mistake in Redner's work has been noticed, the distance in Lemma 4.3 is the first you would think of. Verifying Wald's assumptions does not pose any technical challenge as long as I can see. I keep my score unchanged.

Authorsrebuttal2023-08-16

Thank you for your response and for engaging with us for a productive discussion. We would like to bring the following points to your attention with regard to Redner's work and the verification of Wald's criteria. i) We assume you meant Lemma 4.3. We respectfully disagree that the distance in that result is the first one that comes to anyone's mind and that is because it is not in general a distance on quotient parameter space $X/C$ (it is a distance on some other quotient space $X_{\approx}$). In our case, it happens to be a distance on $X/C$. This connection and why they are the same is far from being trivial. Please refer to our response to the first question of Reviewer gPBN where we exclusively discuss this point and consider looking at the reference provided there. ii) Redner (1981), `Note on the consistency of the maximum likelihood estimate for nonidentifiable distributions', Annals of Statistics is a fairly cited paper in the relevant literature. But if you look deeper into those citations, you will find it is mostly an acknowledgement saying that the situation has been handled within a quotient space (which is, unfortunately, a misconception). The methodologies were never used afterwards as the foundation was shaky. Therefore, it is important to have a correct foundation for it to be useful for future practitioners and clear up certain misconceptions. In our work, we tried our best to maintain neutral language while pointing out these issues. Contrary to Redner's work, our framework is not limited to an abstract piece of mathematics as it allows us to prove consistent covariance estimation as we pointed out in the response of the previous reviewer. Therefore, it is useful for practitioners and more experimentally inclined researchers for implementation purposes. iii) With regard to your last point, it is true that verifying Wald's conditions is far less challenging than actually building the theory in the correct way. With that said it has its challenges. As we said, in general, sup (or inf) over an uncountable family of functions gives rise to measurability issues which are hardly addressed (or even mentioned) in the previous works. Also in the third line of our proof of Lemma 6.3 (in the upper bound of log-likelihood), the short argument relied on the homoskedastic error term assumption for the PPCA model. Had this been a Factor model with heteroskedastic error term the proof would not have been the same. Therefore we tried to exploit the model assumptions on PPCA whenever we could and it is not merely a replication of something more general.

Program Chairsdecision2023-09-21

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC