The Star Geometry of Critic-Based Regularizer Learning

Variational regularization is a classical technique to solve statistical inference tasks and inverse problems, with modern data-driven approaches parameterizing regularizers via deep neural networks showcasing impressive empirical performance. Recent works along these lines learn task-dependent regularizers. This is done by integrating information about the measurements and ground-truth data in an unsupervised, critic-based loss function, where the regularizer attributes low values to likely data and high values to unlikely data. However, there is little theory about the structure of regularizers learned via this process and how it relates to the two data distributions. To make progress on this challenge, we initiate a study of optimizing critic-based loss functions to learn regularizers over a particular family of regularizers: gauges (or Minkowski functionals) of star-shaped bodies. This family contains regularizers that are commonly employed in practice and shares properties with regularizers parameterized by deep neural networks. We specifically investigate critic-based losses derived from variational representations of statistical distances between probability measures. By leveraging tools from star geometry and dual Brunn-Minkowski theory, we illustrate how these losses can be interpreted as dual mixed volumes that depend on the data distribution. This allows us to derive exact expressions for the optimal regularizer in certain cases. Finally, we identify which neural network architectures give rise to such star body gauges and when do such regularizers have favorable properties for optimization. More broadly, this work highlights how the tools of star geometry can aid in understanding the geometry of unsupervised regularizer learning.

Paper

Similar papers

Peer review

Reviewer kSE28/10 · confidence 3/52024-07-11

Summary

The paper presents a theoretical analysis of learning regularizers for inverse problems using a critic-based loss. By focusing on a specific family of regularizers (gauges of star-shaped bodies), amenable to theoretical analysis, the authors provide a number of theoretical insights towards existence, uniqueness within existing frameworks (based on wasserstein distance) and further extensions to f-divergences. This is further connected to the existing literature on learned regularization by considering star bodies corresponding to weakly convex regularisers.

Strengths

The paper presents a very novel idea of utilising the theoretical framework of star-bodies in order to provide theoretical interpretability of critic based regularization. Given the recent interest in learned regularization, this paper opens up a number of new research directions both theoretically and numerically.

Weaknesses

There are a few weaknesses, which in my opinion are not limiting. To be precise, the paper focuses on a specific class of regularizers (gauges of star-shaped bodies) and a specific type of critic-based loss functions (derived from variational representations of statistical distances). It would be interesting to see if the results can be extended to other classes of regularizers and loss functions. The paper primarily investigates this class of regularizers theoretically and with very few experiments. The paper does not include any experiments to demonstrate the practical performance of the learned regularizers, and while the theoretical results are valuable, it would be helpful to see how they translate into practice. This could be of interest as future work for practitioners working on inverse problems.

Questions

The paper is very well-writen, and as such there are very few questions that I have: * One of the motivations, also discussed in 1.2 (and line 57), is that uniqueness of the transport potential does not hold when considering the wasserstein 1 based loss. I would like to refer the authors to arxiv.org/abs/2211.00820, as in fact (under some assumptions), this uniqueness can be shown to be unique $D_n$ almost everywhere. With this result in mind, could you explain intuitively why in Theorem 2.4, it is possible to prove uniqueness without a.e.? * Line 130 "this map" - which map is this referring to? * I am not entirely sure what the relevance of Remark 2.8 is. In practice, rescaling the distribution destroys information from the true distribution - the critic that is desired is the one that would be operating on $D_r$ and $D_n$, and not $D_r$ and $\lambda D_n$. * It would be very intersting to see whether the optimal regularisers derived as minimisers of the variational objective are also optimal regulariser in the sense of Leong et al. *Line 50 "about the measurements". Clasically the measurements themselves live in a different space from the original data. For this reason Lunz et al. utilises backprojection to first map it to the same space. * Line 29 "ill-posed meaning that there are an infinite number ..." - In the inverse problem literature, ill-posedness does not correspond to non-uniqueness only. I suggest referring to Hadamards definition of well-posedness. E.g. see Shumaylov et al. or Arridge, Simon, et al. "Solving inverse problems using data-driven models." Acta Numerica 28 (2019): 1-174.

Rating

8

Confidence

3

Soundness

3

Presentation

4

Contribution

4

Limitations

* See weaknesses.

Reviewer XeHQ6/10 · confidence 2/52024-07-13

Summary

This paper explores the learning of task-dependent regularizers using critic-based loss functions in the context of variational regularization for statistical inference and inverse problems. It particularly focuses on a specific family of regularizers, namely gauges of star-shaped bodies, which are common in practice and similar to those parameterized by deep neural networks. The study introduces a novel approach utilizing tools from star geometry and dual Brunn-Minkowski theory, which allows the derivation of exact expressions for the optimal regularizer under certain conditions and explores the properties of neural network architectures that can yield such regularizers. This work contributes to a deeper understanding of the structure of data-driven regularizers and their optimization characteristics.

Strengths

The problem setup and the theoretical framework of the paper appear rigorous and methodologically sound. The motivation behind the study is robust, addressing the theoretical gaps in understanding how regularizers are learned.

Weaknesses

Some of the results presented in the paper are complex and difficult to interpret, which may limit their accessibility to a broader audience. Moreover, the paper does not clearly articulate the practical implications of these theoretical findings for real-world applications, which could hinder its impact.

Questions

1. Could you provide additional context on the adversarial regularization framework, particularly regarding the roles and definitions of the two distributions $D_r$ and $D_n$? 2. The significance of the results in Theorem 3.1 is not clear to me. Could you elaborate on why these results are important and what they contribute to the field of regularizer learning? 3. The paper lacks experimental validation of its theoretical constructs through simulations or empirical data. Some experiments (even simple ones) will be helpful for better understanding.

Rating

6

Confidence

2

Soundness

3

Presentation

2

Contribution

3

Limitations

Yes.

Reviewer 3HXp5/10 · confidence 4/52024-07-13

Summary

This paper leverages the star geometry and dual Brunn-Minkowski theory to study the optimal critic-based regularizers. The authors illustrate the optimal regularizer can be interpreted using dual mixed volumes that depend on data distribution. Theorems are proved for the existence and uniqueness of the optimal regularizer. The authors also identify the neural network architectures for learning the star body gauges for the optimal regularizer.

Strengths

The paper leverages the star geometry in understanding the geometry of unsupervised regularizer learning.

Weaknesses

As cited in the submission, this paper is closely related to [50]. Many concepts, theoretic results, even examples resemble or coincide with those in [50]. The paper failed to clearly distinguish itself from the existing work [50].

Questions

In what scope the submitted work extends [50]? On line 68, "assigns" should be "assigned"? On line 118, what does "[x, y]" mean if both x, y are points in R^d? Did you mean the line segment connecting x and y?

Rating

5

Confidence

4

Soundness

3

Presentation

3

Contribution

1

Limitations

See the weakness

Reviewer 5WpE6/10 · confidence 3/52024-07-15

Summary

This submission extends the techniques of [50], i.e., tools from star geometry and dual Brunn-Minkowski theory, to characterize the optimal regularizer under the adversarial regularization framework [51] of inverse problems. \alpha-Divergence as loss functions for learning regularizers is also discussed, with the dual mixed volume interpretations. The weak convexity and compatible neural network architectures are further discussed for computational concerns related to the proposed star body regularizers.

Strengths

Extending the analysis and results of [50] to the adversarial regularization framework and showing its connections to the \alpha-divergence is an interesting theoretical contribution. The specified neural network layers compatible with the star body regularizer can also shed light on practice.

Weaknesses

My major concern with this work is that it is unclear if the proposed new \alpha-divergence-based loss functions are useful for the adversarial regularization problem this submission studies. The original adversarial regularization work [51] for inverse problems, albeit published in NeurIPS 6 years ago, has reported experimental results to validate the proposed framework. On the other hand, if positioned as a pure theory work, given the existence of [50], I feel that the theoretical contribution of this submission seems a bit short for publication in NeurIPS. As a minor thing, it would be helpful for readability if an overview of the organization and the flow of the paper could be briefly presented in the introduction. After the authors' adding new experiments during the discussion phase --------------------------------------------------------------------------------------------- These two concerns are partially addressed. On one hand, I can see the potential of the proposed approach from these experiments. On the other hand, the experiments are still preliminary and small scale.

Questions

Prop. 4.3 requires positive homogeneity of each layer, which limits the choice of activation functions in practice to a subset of piecewise linear functions, such as ReLU and its variants. I feel that this could be a limitation of the proposed approach. Moreover, while each layer is an injective function, due to the homogeneity of activations, the composite of such layers will not be an injective (see Sec. 3 of the paper below for example), will this break the proof of Prop. 4.3? Dinh, Laurent, et al. "Sharp minima can generalize for deep nets." ICML 2017. After the authors' rebuttal --------------------------------------------------------------------------------------------------- The question regarding the composite of injectivities does not apply to Prop. 4.3.

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

Mostly. Please refer to the Questions session for a concern I have regarding limitations.

Reviewer 3HXp2024-08-08

Thanks to the authors for their clarification. I see the contribution better and will raise my score.

Authorsrebuttal2024-08-09

Thank you to the reviewer for taking the time to consider our rebuttal. We sincerely appreciate raising your score.

Reviewer kSE22024-08-09

Thanks to the authors for their clarifications and answers. I would suggest adding a note about the 'almost everywhere' as a distinguishing factor from the rest of the literature. I would be very happy to see the paper accepted.

Authorsrebuttal2024-08-09

Thank you to the reviewer for taking the time to consider our rebuttal and for the suggestion. We agree that this is a good point to highlight, and will be sure to do so in a revised version of the manuscript. We sincerely appreciate the reviewer's support of our paper.

Reviewer 5WpE2024-08-09

Thanks to the authors for the detailed response to concerns and questions, which helps me to better access the novelty and theoretical contribution of this submission. As a result, I have raised my rating. My previous question regarding the composite of injectivities actually apply to the mappings from parameters to network functions, rather than the input-output network functions, so this does not apply to Prop. 4.3 and the authors' explanation is correct. With all that said, my concerns regarding limited activation function choices and lack of experimental validations remain. That is why I am not able to raise the rating to a firm accept for NeurIPS.

Authorsrebuttal2024-08-12

Thank you to the reviewer for considering our rebuttal and for raising their rating. Based on your concerns, we conducted experiments on two points that were raised, namely the applicability of the $\alpha$-divergence based loss and the importance of the homogeneity of the activation function. Please see the official comment posted titled "New Experimental Results". Thank you for voicing your concerns and suggesting we explore this further. We believe that these additional experiments are valuable in providing support for our theory and will significantly improve the paper quality. Please let us know if you have any additional questions or comments. We hope that these experiments can potentially address your concerns.

Reviewer 5WpE2024-08-12

I would like to thank the authors for conducting experiments regarding my two major concerns. While the experiments are quite preliminary and small-scale, I can see the potential of the proposed approach from them. Therefore, I will further raise my rating.

Authorsrebuttal2024-08-13

Thank you to the reviewer for appreciating our experiments and acknowledging the practical potential of our theory. We sincerely appreciate raising your score.

Authorsrebuttal2024-08-12

New Experimental Results

Dear reviewers, We sincerely thank you for your time spent reviewing our work and responding to our rebuttal. We wanted to update you all to let you know that we have conducted experiments based on the comments from some reviewers. In particular, some reviewers expressed concern regarding our lack of empirical results to support our theory. To address these concerns, we have conducted two experiments: one on using the new $\alpha$-divergence loss to learn regularizers and the other on the requirement of positive homogeneity in the neural network layers. **Denoising comparison of Hellinger and Adversarial Regularizers:** We wanted to test the performance of regularizers learned using the Hellinger-based loss of Eq (5) and those learned using the adversarial loss in Theorem 2.4. To do this, we consider denoising on the MNIST dataset. We take $1000$ random samples from the MNIST training set (constituting our $\\mathcal{D}_r$ distribution) and add Gaussian noise with variance $\\sigma^2 = 0.05$ (constituting our $\\mathcal{D}_n$ distribution). We then aim to reconstruct *test* samples from the MNIST dataset corrupted with Gaussian noise of the same variance seen during training. For our regularizers, we parameterized them with a deep convolutional neural network, similar to the construction in Lunz et al [51]. Specifically, the network has 8 convolutional layers with LeakyReLU activation functions and an additional 3-layer MLP with LeakyReLU activations and no bias terms. The final layer is the Euclidean $\\ell_2$ norm. This network implements a star-shaped regularizer, as outlined in the discussion of our Theorem 4.3. The regularizers were trained using the adversarial loss and Hellinger-based loss (Eq (5)). We also used the gradient penalty term from Lunz et al [51] for both losses. After training, we then aim to reconstruct noisy *test* samples $y = x_0 + z$ where $x_0$ is a test digit and $z \sim \mathcal{N}(0,\sigma^2I)$. We do this by solving the following with gradient descent initialized at $y$: $$\\min_{x} ||x - y||^2 + \\lambda \mathcal{R}(x).$$ Here, $\\mathcal{R}$ denotes our trained regularizer. For the adversarially trained network, we used $\lambda := 2 \cdot \tilde{\lambda}$ where $\tilde{\lambda} = \mathbb{E}_{z\sim\mathcal{N}(0,\sigma^2I)}||z||_2$, as described in [51]. For the Hellinger-based network, we found that $\\lambda = 10 \tilde{\lambda}^2$ gave better performance, so we used this for recovery. We display the average PSNR and MSE over 20 test images for both networks below. We see that the Hellinger-based loss gives competitive performance as compared to the adversarially trained network. This suggests that these new $\alpha$-divergence based losses are potentially worth exploring from a practical perspective as well. **Noisy image MSE, PSNR:** 0.0486, 13.13 **Hellinger MSE, PSNR:** 0.0046, 23.52 **Adversarial MSE, PSNR:** 0.005, 23.14 **The role of the activation function:** While our theory for neural networks is mainly limited to activation functions such as ReLU and LeakyReLU due to their positive homogeneity, we wanted to understand whether this was a limitation or beneficial from an empirical perspective. Hence we analyzed the influence of the activation function in the above network's performance in denoising. For this study, we changed all LeakyReLU activations in the network to be either the Exponential Linear Unit (ELU), the Gaussian Error Linear Unit (GELU), or the Tanh activation. We trained these networks using the adversarial loss as previously described and then used them for denoising on the same images from the previous experiment. As compared to the network with LeakyReLU activations, we see significant performance degradations when switching to these non-positively homogenous activations. This suggests that potentially positive homogeneity of the activations is a useful property in terms of performance of the neural network-based regularizer. **LeakyReLU MSE, PSNR:** 0.005, 23.14 **ELU MSE, PSNR:** 0.042, 14.16 **GELU MSE, PSNR:** 0.031, 15.29 **Tanh MSE, PSNR:** 0.049, 13.14 We sincerely thank the reviewers again for their engagement and for raising these concerns. We believe that these additional experiments are valuable in providing support for our theory and will significantly improve the paper quality. We will be sure to include these in the updated version of manuscript. Please let us know if you have any additional questions or concerns.

Reviewer kSE22024-08-13

I would like to thank the authors for the even more extensive numerical comparison. While I would not consider MNIST to be a good example dataset, it is nonetheless more informative than currently provided experiments. I would like to emphasise two particular points which are a bit problematic in the comparison above. In my understanding, in [51] $\lambda$ can be chosen in the form provided, as the regulariser is assumed to be 1-lipshitz. However, by turning to star-bodies I do not believe you have that property, right? This certainly becomes more problematic when considering Hellinger based distances, as you have seen in practice - you need a very different regularisation parameter. As such, above seems like a very fine-tuned comparison, not actually representative of performance, as the Adversarial MSE can likely be fine-tuned to be even better performing by tuning that hyperparameter. I would also suggest including a TV or l1 regulariser to show comparison to model-based regularisation as some sort of baseline. All in all however, I find this very encouraging, and if the authors agree with the points above and are happy to add these experiments to the paper with a discussion of the points above regarding parameter choice, I would be happy to increase my score.

Authorsrebuttal2024-08-13

Response to Reviewer kSE2

We thank the reviewer for appreciating the new experimental results and for bringing up these interesting points. Below we further delve into the two main points: **Lipschitzness and regularization strength:** The reviewer is correct that, in general, star body/star-shaped gauges may not be 1-Lipschitz. There are certain conditions that can guarantee Lipschitzness. For example, in the paper, we mention a result that for a star body $K$, the map $x \mapsto ||x||_K$ is 1-Lipschitz if and only if the kernel of $K$ contains the unit Euclidean ball $B^d \subseteq \mathrm{ker}(K)$. When $K$ is star-shaped and not a star body, this condition no longer becomes necessary and sufficient, but there are other ways one can achieve Lipschitzness. For example, when the star-shaped regularizer is parameterized by a deep neural network with positively homogenous activations and linear layers (convolutional or fully connected with no bias terms), it is possible for the regularizer to be 1-Lipschitz by either enforcing each activation to be 1-Lipschitz with linear layers of spectral norm at most 1 or enforcing unit gradient norm, as done in [51]. We agree with the reviewer's intuition that the Hellinger-based loss potentially learns a network that is less Lipschitz than the one learned via Adversarial regularization, hence resulting in the need for a different choice of regularization strength when used for inverse problems. In addition to the reviewer's point, we hypothesize this may also result from the fact that the Hellinger loss and adversarial loss weight the distributions $\mathcal{D}_r$ and $\mathcal{D}_n$ differently, resulting in different regularity properties of the learned regularizer. We mainly fixed the Adversarial regularizer's $\lambda$ to the reported value because this was the value recommended in [51], but we will be sure to perform further testing on $\lambda$ and add a more in-depth discussion regarding how these choices relate to Lipschitzness. **Including baselines:** Thank you for this suggestion. We went with the reviewer's suggestion and considered TV regularization. Under the same experimental setup as before, after tuning the regularization strength, we found that TV yielded the following MSE and PSNR: **TV MSE, PSNR:** 0.009, 20.3 To summarize, we appreciate and fully agree with the reviewer's suggestions on providing a more in-depth discussion regarding hyperparameter choices. In an updated version of the manuscript, we will add - a discussion regarding parameter choices for both methods when employed to solve inverse problems - updated results with further testing on parameter choices for all methods on more images - TV denoising results as a model-based baseline We thank the reviewer again for their continued support and encouragement. We sincerely appreciate raising your score!

Authorsrebuttal2024-08-12

Dear reviewer, Thank you again for your detailed review and positive comments. We wanted to provide you with a brief update on our work. Based on your and other reviewer's comments, we have conducted new experiments to support our theory. Please see the official comment titled "New Experimental Results". We hope that these experiments along with our rebuttal can potentially address your concerns.

Authorsrebuttal2024-08-13

We thank the reviewer for considering our rebuttal and keeping their positive score.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC