Matrix Denoising with Doubly Heteroscedastic Noise: Fundamental Limits and Optimal Spectral Methods

We study the matrix denoising problem of estimating the singular vectors of a rank-$1$ signal corrupted by noise with both column and row correlations. Existing works are either unable to pinpoint the exact asymptotic estimation error or, when they do so, the resulting approaches (e.g., based on whitening or singular value shrinkage) remain vastly suboptimal. On top of this, most of the literature has focused on the special case of estimating the left singular vector of the signal when the noise only possesses row correlation (one-sided heteroscedasticity). In contrast, our work establishes the information-theoretic and algorithmic limits of matrix denoising with doubly heteroscedastic noise. We characterize the exact asymptotic minimum mean square error, and design a novel spectral estimator with rigorous optimality guarantees: under a technical condition, it attains positive correlation with the signals whenever information-theoretically possible and, for one-sided heteroscedasticity, it also achieves the Bayes-optimal error. Numerical experiments demonstrate the significant advantage of our theoretically principled method with the state of the art. The proofs draw connections with statistical physics and approximate message passing, departing drastically from standard random matrix theory techniques.

Paper

Similar papers

Peer review

Reviewer Fvn77/10 · confidence 3/52024-06-30

Summary

This paper studies the following "matrix denoising"/low-rank estimation problem. Given a rectangular matrix $Y = uv^\top + W$, the goal is to recover the rank-one factors $u,v$, when $W$ is random, with as little $\ell_2$ error as possible. A classic and well-studied setting takes $W$ to have iid entries; often one even assumes that the entries are Gaussian. This setting has been studied in random matrix theory, high-dimensional statistics, and more recently via techniques from statistical physics. The iid noise assumption is quite strong, and so a more recent line of works aims to relax that assumption by allowing correlations among the entries of $W$. In the "two-sided" heteroscedastic noise model, $W$ takes the form $\Theta^{1/2} \cdot W' \cdot \Sigma^{1/2}$, where $\Theta$ and $\Sigma$ are PSD matrices and $W'$ has iid Gaussian entries. The assumption is that $\Theta$ and $\Sigma$ are known. The main contributions of this paper are twofold: - a new spectral estimator whose performance improves over more naive spectral methods for recovery of $u$ and $v$. - a proof that under slightly stronger assumptions ($u$,$v$ Gaussian), the spectral estimator obtains nontrivial $\ell_2$ error whenever this is information-theoretically possible, and even obtains information-theoretically optimal $\ell_2$ error when the heteroscedasticity is one-sided. Along the way, the paper establishes a simple formula for the information-theoretic signal-to-noise threshold governing when nontrivial recovery of $u$ and $v$ is possible in this model. The paper also performs numerical experiments on synthetic data to validate the theory.

Strengths

- Well-written and easy to read first 9 pages - Thorough mathematical investigation of the two-sided heteroscedastic spiked matrix model - New algorithm with nice optimality guarantees

Weaknesses

My main reservation is that, as with a lot of papers using stat phys techniques, one gets the feeling that the assumptions are sort of designed to make the mathematical techniques work. I think the biggest offender here is the assumption that $\Sigma$ and $\Theta$ are known. Less major are the Gaussian-ness assumptions, the assumption that $n,d$ are of comparable order, and the assumption that the empirical spectral distributions of $\Theta$ and $\Sigma$ converge -- I think these allow the use of asymptotic methods, basically.

Questions

none

Rating

7

Confidence

3

Soundness

4

Presentation

4

Contribution

3

Limitations

Yes

Reviewer hDgq7/10 · confidence 2/52024-07-12

Summary

The authors study how to recover a rank one spike corrupted by doubly heteroscedastic gaussian noise in the high dimensional regime. We are given a condition on the signal to noise ratio to indicate whether it's information-theoretically possible to have a non-trivial recovery of the spike. If this is satisfied then there is spectral estimator which can obtain non-trivial recovery. In particular cases this estimator is also Bayes optimal.

Strengths

The paper is well written and does a really good job at introducing the main ideas in an intuitive way. While this kind of spectral estimators are well known in the physics literature, their application to such a noise model is an interesting generalisation. I believe this paper thoroughly explores this denoising problem, with the only easily achievable extension being looking at a rank $r$ spike in the signal (where r is a constant in d,n), which I am sure however wouldn't alter the results significantly. Thus this is in my opinion quite a solid contribution with not much room for improvement.

Weaknesses

The paper introduces no significant advances in the theoretical tools or understanding of matrix denoising as all the tools used are essentially well known. I believe one issue in the writing is its lack of clarity in stating which results are completely rigorous and which aren't. The authors are upfront in saying they are guided by physics-inspired heuristics, but reading in the appendix it seems like (at least in some sections) the derivations are quite solid. A tangible improvement to the writing would be to state explicitly which results are conjectures and which are theorems. I think it would greatly improve the readability of figures 2 and 3 to have the error on the mean instead of the std. I would add a sentence describing more explicitly the difference and respective advantages of (3.1) (where the spectrum is O(1)) and (4.1) (where the elements of Y are O(1)). Small typos: 1. Line 203: $\sigma_2$ is not directly defined. You will only do so in Theorem 5.1 2. Line 188: you invoke the Nishimori identity but don't state it in the main text 3. Line 200: you say that $\eta>0$. While I also expect this to be true, I think you mean to say that they are real. The same applies to all the square roots in 5.3 and 5.4 Finally, not exactly a typo but I personally find it confusing to use $\bar \Sigma$ and $\bar \Xi$ when one wants to average over the singular values of $\Sigma$ and $\Xi$. I would prefer having an explicit integral over the singular value PDF.

Questions

Would it be possible to look at the singular values of $A$ before and after the pre-processing? Can we clearly see a spike emerging if 5.1 is true by doing the pre-processing? You state that AMP has the fundamental limitation that requires a "warm start" to be effective. While this is true, initialising the estimator at random from the prior should allow it to have non-zero overlap in O(log(d)) steps. Do you see this in your numerics? Could you run AMP after using the spectral estimator and put the additional lines and simulation dots in figures 2, 3?

Rating

7

Confidence

2

Soundness

3

Presentation

3

Contribution

2

Limitations

The limitations are correctly addressed in section 6. There is no negative societal impact of this work.

Reviewer S4fP8/10 · confidence 1/52024-07-13

Summary

This paper considers the problem of matrix denoising. Given an observation X = A + W where W is noise, our goal is to estimate A, which is typically low-rank. Unlike previous works, this paper treats the case where W is doubly heteroscedastic. The authors identify a condition for non-trivial estimation of the signal vectors, along with an accompanying spectral algorithm that succeeds whenever that condition holds, under a technical condition.

Strengths

The paper addresses an important problem using a novel approach using statistical physics/ AMP concepts. The results, both theoretical and empirical, are strong and add considerably to the literature.

Weaknesses

Strictly speaking, you do not show that whitening fails, only that the whitened matrix does not match the proposed AMP approach (lines 262-265).

Questions

Initially you say that you believe (5.1) implies \sigma_2^{\star} < 1, and later you say that you believe these conditions are equivalent. Could you please clarify which one it is?

Rating

8

Confidence

1

Soundness

4

Presentation

4

Contribution

4

Limitations

Yes

Reviewer hDgq2024-08-13

Thanks for your detailed rebuttal, I will raise my score accordingly. For the figure, I am slightly surprised the fluctuations are so large, I guess trying larger sizes might be doable and make the presentation clearer.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC