Summary
Thanks for the author's rebuttal. I have read all of the rebuttals and reviews and decided to keep the current rating.
Current work proposed an extension of the variational diffusion model, called the variational latent diffusion model. It is applied to inverse problems in the high energy physics field and tested on a single topology. The result shows that the distance is three times less than current latent diffusion models.
Strengths
Originality: The model is a combination of two existing models, the latent diffusion model and the variational diffusion model. Moreover, the author defined an appropriate loss function to train the model.
Quality: The result demonstrated that the proposed model outperforms other baseline models.
Significance: The experiment provides unique data in the high energy physics field.
Weaknesses
The paper has several typos and unclear points. Moreover, some claims are not well supported by the experiment results. The reviewer is confused to some presented points. Details can be found in the questions part.
Questions
1. There is no general caption for Figure 1. Moreover, the sub-caption is not consistent with the sub-figure. For example, ($e$) in caption while ($e^{-}$) in the subfigure. ($\nu$) in caption while $\bar{v}$ in subfigure. $d$ in caption while $\bar{d}$ in subfigure. Moreover, subfigure (b) is too small to see the detail; the reviewer did not get the point of how it is related to your problem or the model. The font in this subfigure is too small to see. The review suggests presenting the subfigure (b) in a better way to give more insights into your problem or model.
2. There are several typos in the paper. For example, line 8, ``latent variation diffusion model``, is supposed to be ``latent variational diffusion models`` Line 192, the $\hat{\cdot}$ is on $z_t$ or $t$?
3. Some definitions are not clear. For example, how did you calculate $\hat{E}$ and $\hat{p}$ from the predicted value in Equation 10? What is the norm of the $\lambda_c|\cdot|$ in Equation 10? Is it $L_1$ norm or $L_2$ norm?
4. The author uses a deterministic encoder and claims it is better than the variational encoder. According to the author, the latter only has limited benefits while increasing training variance and complexity. However, there is no experiment result to support the claim. More ablation study is needed to justify the claim. Still, it sounds strange to the reviewer that the model is called a variational latent diffusion model but only uses a deterministic encoder.
5. The author proposed three variants of the model. Conditional, unconditional, and the last one, conditional encoder and unconditional decoder. The way of presenting the results confuses the reviewer. Firstly, which one do you promote to use in the conclusion? It seems in Table 1. VLD is better in the latter three metrics, while the UC-VLD is slightly better in the first three metrics. However, Figure 2 shows the framework of VLD, while all the Figures in the result part and the appendix part only show the results for the UC-VLD. If the authors want to promote VLD, consider adding relevant plot results for VLD. If the UC-VLD is better, consider changing Figure 2 to reflect this point.
6. A physics-informed consistency loss is proposed with a hyperparameter $\lambda_{c}$; how did you adjust the value of this hyperparameter? Moreover, what would be the effect of including and not including this loss?
7. The author claims the model is aimed at high-dimensional inverse problems. However, the training cost for $55$ dimensional variables is expensive, $24$ hours. Moreover, the designed latent space has a higher dimension than the input instead of the commonly used lower latent space dimension for compression. This design could introduce additional costs. And the reviewer is concerned with the scalability of the current model to higher dimensional problems.
8. What is the limitation of the current work?
Rating
4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Limitations
The author did not mention the limitations explicitly. I would be curious what would be the limitation of the proposed model.