Looks Too Good To Be True: An Information-Theoretic Analysis of Hallucinations in Generative Restoration Models

The pursuit of high perceptual quality in image restoration has driven the development of revolutionary generative models, capable of producing results often visually indistinguishable from real data. However, as their perceptual quality continues to improve, these models also exhibit a growing tendency to generate hallucinations - realistic-looking details that do not exist in the ground truth images. Hallucinations in these models create uncertainty about their reliability, raising major concerns about their practical application. This paper investigates this phenomenon through the lens of information theory, revealing a fundamental tradeoff between uncertainty and perception. We rigorously analyze the relationship between these two factors, proving that the global minimal uncertainty in generative models grows in tandem with perception. In particular, we define the inherent uncertainty of the restoration problem and show that attaining perfect perceptual quality entails at least twice this uncertainty. Additionally, we establish a relation between distortion, uncertainty and perception, through which we prove the aforementioned uncertainly-perception tradeoff induces the well-known perception-distortion tradeoff. We demonstrate our theoretical findings through experiments with super-resolution and inpainting algorithms. This work uncovers fundamental limitations of generative models in achieving both high perceptual quality and reliable predictions for image restoration. Thus, we aim to raise awareness among practitioners about this inherent tradeoff, empowering them to make informed decisions and potentially prioritize safety over perceptual performance.

Paper

Similar papers

Peer review

Reviewer DvxH7/10 · confidence 3/52024-06-21

Summary

This work aims to provide a theoretical analysis about the uncertainty-perception trade-off in generative models, corresponding to the fidelity-naturalness trade-off of the generated images. By defining the inherent uncertainty and formulating a uncertainty-perception (UP) function, the authors proves that the UP function is globally lower-bounded by the inherenet uncertainty. Additionally, they derive that perfect perceptual quality requires at least twice the inherent uncertainty. The proposed theoretical framework establish a relationship between uncertainty and MSE, resembling the well-known perception-distortion trade-off. All theoretical findings are empirically verified with image super-resolution algorithms.

Strengths

+ This is a timely theoretical analysis about the hallucination phenomenon widely occurs in generative models. + The proposed theoretical framework helps practitioners better understand the tradeoff between uncertainty and perceptual quality, guiding them to tune the models in real-world safety-sensitive applications. + Detailed proofs are provided.

Weaknesses

- In my opinion, LPIPS, MSE, PSNR, SSIM are all full-reference image quality metrics that quantify the fidelity of the restorted images with respect to the ground-truths. Would it be more reasonable to adopt no-reference image quality metrics to quantify the perception ? - The experiment part is relatively weak, where more quantitative examples are expected. - The GT image is not presented in Fig. 5, making it difficult to assess the fidelity of the resorted results.

Questions

I think it would be more convincing if the authors can provide more experimental results in real-world SR tasks.

Rating

7

Confidence

3

Soundness

3

Presentation

4

Contribution

3

Limitations

Yes.

Reviewer WuZt6/10 · confidence 4/52024-07-08

Summary

This paper presents a theoretical perspective towards hallucinations and reveals a tradeoff between uncertainty and perception for image restoration problem. Additionally, the paper points out that uncertainly-perception tradeoff can induce the well-known perception-distortion tradeoff.

Strengths

1. The paper provides a theoretical interpretation about the hallucinations problems of inverse problem, which may offer useful guidance for further practical research. 2. The writing is good and the paper is easy to follow. 3. Rich theoretical results about uncertainly-perception tradeoff and its relationship with perception-distortion tradeoff.

Weaknesses

1. Typically, perception can be measured by criteria like LPIPS. Why can the conditional convergence (Eq 1) provide measurement for perception? Can the authors provide some intuitive explanation? 2. The theorem 2 is based on a strict assumption that $D_v$ is convex in its second argument. Can you proof this directly? 3. The visual results are not enough. The authors should provide some examples that contain hallucinations. It seems that in Figure 5, the details are realistic-looking details but not hallucinations.

Questions

see weakness

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The authors discuss the limitations.

Reviewer WuZt2024-08-09

After Rebuttal

Thanks for the rebuttal. The authors address my questions. I strongly recommend the authors to add more visual results to illustrate the paper clearly. I raise the score to 6.

Authorsrebuttal2024-08-11

We deeply value the reviewer's positive feedback and the subsequent score increase.

Reviewer xyrt6/10 · confidence 3/52024-07-11

Summary

The paper employs information-theory tools to characterize a tradeoff between uncertainty and perception in image restoration. They prove that high perceptual quality leads to increased uncertainty and the uncertainty-perception trade-off induces the distortion-perception trade-off. The theoretical results are illustrated with experiments in image super-resolution tasks.

Strengths

1. The paper is well-written. 2. The authors provide clear theoretical and quantitative demonstrations.

Weaknesses

1. The trade-off between perception and distortion has been discussed in some previous works. The authors should clarify their contribution especially compare to previous works[1,2]. 2. This paper establishes the theoretical relationship between uncertainty and perception. However, the authors do not provide practical applications, e.g. how to use this relationship in restoration task. Refs: [1] The Perception-Distortion Tradeoff. CVPR, 2018. [2] The Perception-Robustness Tradeoff in Deterministic Image Restoration. arXiv:2311.09253.

Questions

1. The uncertainty-perception plane is based on SFID, PDL and LPIPS, the relationship is still exists on stronger vision-language IQA like LIQE[1], Q-ALIGN[2]? 2. Some recent SR methods[3,4] explore a better trade-off between perception and artifacts. How they perform in uncertainty-perception and uncertainty-distortion measurement? Refs: [1] Blind image quality assessment via vision-language correspondence: A multitask learning perspective. CVPR, 2023. [2] Q-ALIGN: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels. arXiv:2312.17090. [3] DeSRA: Detect and Delete the Artifacts of GAN-based Real-World Super-Resolution Models. ICML, 2023. [4] Details or artifacts: A locally discriminative learning approach to realistic image super resolution. CVPR, 2022.

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

2

Limitations

See weakness

Reviewer W9yk6/10 · confidence 4/52024-07-12

Summary

Deep generative models have achieved remarkable performance in image restoration, resulting in generated images of high visual quality. However, these models often produce high-frequency details that are not consistent with the ground-truth images. Such hallucinations introduce uncertainty in the generated content and affect the reliability of model predictions. This paper defines the uncertainty of image restoration and the uncertainty-perception (UP) function, and reveals an uncertainly-perceptual trade-off. The paper theoretically analysis the relationship between the uncertainly-perceptual trade-off and the perceptual-distortion trade-off. The theoretical findings are validated through experiments with image super-resolution algorithms. Results show that no model can achieve both low uncertainty and high perceptual quality simultaneously.

Strengths

1. This is a timely paper that analyses the hallucination of deep generative models. It proposes a novel insight of the phenomenon of hallucinations in generative models, a critical issue that affects the reliability of image restoration tasks. 2. The paper adopts the Bayesian framework to analyze the tradeoff between uncertainty and perception. This framework helps in quantifying the inherent uncertainty in generative models and establishes a theoretical foundation for understanding the limitations of these models. 3. The theoretical findings are empirically validated using single-image super-resolution algorithms. This strengthens the credibility of the theoretical analysis of the study. 4. Following the definition, a concrete example (e.g., example 1) is illustrated which can help the readers better understand the concept.

Weaknesses

1. Experiments only provide results in image super-resolution, how about applying the proposed method on other restoration tasks? 2. The experiments primarily use synthesized datasets, such as BSD100. The paper would benefit from including experiments on more diverse and real-world datasets to validate the findings under different conditions and data distributions. 3. The paper relies on entropy and Rényi divergence for theoretical analysis. However, the practical estimation of high-dimensional entropy is challenging. Although the authors use a tractable upper bound for uncertainty, the practical estimation methods and their limitations are not thoroughly discussed

Questions

1. The paper [1] quantifies the structural uncertainty for image restoration. The intuitive explanation is better for readers to understand the inherent uncertainty mentioned in the paper. 2. The practical estimation for computing divergence and uncertainty should be elaborated, which is challenging for real-world images that are typically high-dimensional. 3. While the theoretical analysis is robust, the paper could benefit from more concrete examples of how the findings apply to real-world scenarios, like healthcare or autonomous systems. [1] Belhasin O, Romano Y, Freedman D, Rivlin E, Elad M. Principal uncertainty quantification with spatial correlation for image restoration problems. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2023 Dec 14.

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The paper aims to quantify the potential limitation (e.g., hallucination) of generative models. At the end of the paper, it discusses the limitations of the proposed method.

Reviewer DvxH2024-08-08

Post-rebuttal

Thanks for the reponses, I raised my rating to 7.

Authorsrebuttal2024-08-11

We sincerely appreciate the reviewer recognizing our contribution and raising the score accordingly.

Reviewer W9yk2024-08-11

Thank you for your responses and comprehensive explanation. Your responses have addressed mu concerns, and I will decide to raise the score to WA.

Authorsrebuttal2024-08-11

We are thankful for the reviewer's insightful comments and the resulting score change.

Reviewer xyrt2024-08-12

After Rebuttal

Thank you for the detailed response. I'll keep my score.

Authorsrebuttal2024-08-12

We sincerely thank the reviewer for their valuable feedback, particularly the request to incorporate non-reference perceptual measures, which has significantly strengthened our contribution.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC