Summary
This paper introduces a benchmark for evaluating the out-of-distribution (OOD) performance of neural image compression (NIC) models. The authors observe distinguished spectral biases of NIC models by comparing them with JPEG2000. Then the authors listed all the spectral biases along with a theoretical analysis from the perspective of principle component decomposition.
Weaknesses
I believe the spectral perspective is important to understand NIC models. However, the theoretical analysis is of limited novelty and the analysis and conclusions provided in this paper are empirical and need to be further improved before publication. Moreover, I am afraid that classifications on several corruptions may conflict with the common sense in signal processing field. My concerns are listed as below.
1. The proposed method mainly compares with JPEG2000, which may be inappropriate. JPEG2000 is based on discrete wavelet transform. However, EBCOT and rate control of JPEG2000 could hinder an explicit understanding of the spectral characterization of JPEG2000. I suggest the authors to conduct the comparison experiments on JPEG (which is based on block-based DCT and simple quantization table) for clear comparison. Compared to JPEG2000, JPEG has been widely applied in practice.
2. The comparison presented in Figures 16 and 17 (i.e., comparison to VTM provided in supplementary material) is confusing. $\mathcal{D}$ for VTM seems to be zero but it is a lossy compression method and the reconstruction is still far from lossless (i.e., 37.42 dB).
3. I also have problem with the evaluation using power spectral density (SPD). I noticed that the authors classify Gaussian noise as a high-frequency corruption in Figure 1. However, it is well known that the Gaussian noise, i.e., white noise, has a constant SPD. In other words, the effect of Gaussian noise is expected to be constant rather than constrained on the high-frequency components. Moreover, the effect of JPEG compression is also expected to be on both median and high-frequency, as the default quantization map results in a larger quantization step on the high-frequency components. I am afraid that the classification of other corruptions may neither be sufficiently precise. I guess the gaps between above theoretical analysis and the empirical analysis in Figure 1 originate from the intensity difference on the SPD used to analyze the effect of introduced corruption. The natural images themselves are with an unbalanced SPD. Thus, the observation may be biased.
4. The theoretical analysis in Section 5 is not novel enough. The desired characteristics for transform coding are discussed in [R1]. The differences and similarity between DFT, DCT, and KLT are analyzed in [R2]. The provided analysis on linear transforms (autoencoder) may not be appropriate for nonlinear models [R3].
5. Considering that the major contribution of this paper lies on the empirical evaluations, it is important to make thorough validations on diverse datasets and NIC models. The adopted scale hyperprior is not sufficient supporting the claims in this paper.
[R1] V. K. Goyal, "Theoretical foundations of transform coding." IEEE Signal Processing Magazine, vol. 18, no. 5, pp. 9-21, 2001.
[R2] D. S. Taubman and M. W. Marcellin, JPEG2000: Image compression fundamentals, standards and practice. 2002.
[R3] J. Ballé, et al. "Nonlinear transform coding." IEEE Journal of Selected Topics in Signal Processing, vol. 15, no. 2, pp. 339-353, 2021.
Questions
1. Could the authors offer a comparison with JPEG compressed images? I believe it would be more straightforward.
2. What is the reason for theoretical analysis (given in comment #3 in the section of Weakness) and the given classification (Figure 1) on Gaussian noise and JPEG compression artifacts?
3. What is the difference between the theories for nonlinear transforms (Section 5) and those for traditional linear transform (such as DWT, DCT, and KLT)?
Rating
3: Reject: For instance, a paper with technical flaws, weak evaluation, inadequate reproducibility and incompletely addressed ethical considerations.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.