I would like to thank the reviewers for their rebuttal to address my concerns. I agree with the responses to some of the minor concerns including ablation on roughness, and evaluation settings, but would insist to differ on the rest of the major issues. Specifically,
[1] Unrealistic relit images from the hot dog scene in Fig. 4. The trained image diffusion model is supposed to make correct estimation on the material of the scene (and produce convincing images under novel lighting with correct materials true to the input images). However, assuming the input images are typical of the popular hot dog scene where the bun is mostly diffuse, some samples are unreasonably specular on those regions. I do agree with the author that there is inherent ambiguity between materials and lighting in inverse rendering, which sometimes lead to different plausible combinations of materials and lighting. However, (a) the diffusion model is supposed to make good estimation on the materials based on the input images, instead of producing results of diverse levels of specularity; (b) given multiview observation of the hot dog scene as input, where specularity is NEVER observed on any regions of the bun, there is minimal ambiguity between materials and lighting. There is little to no chance that the bun is actually of specular material while none of the multi-view input images suggest any specularity on the bun. In general, the ambiguity is more likely with single or few images, and more often in the case of baked shadows in materials, and less or not likely at all when there is sufficient observation (as is the setting with this paper) and no specularity is present in any view. In conclusion, I would incline to insist that the samples where the bun is specular are not plausible given the input images, and the trained diffusion model does not do a great job in estimating the true materials in some cases.
Interestingly, in the additional 3 samples of the hot dog scene in the rebuttal PDF, the bun is not specular, while the bun is specular in sample #1, 2, 4, 5, of the 5 samples in Fig. 4(a). It is interesting to hear more from the authors on this difference to make a more informed conclusion on the behavior of the diffusion model.
[2] On the benchmark with Standford-ORB. The authors included additional discussion on the capturing setting of the lighting GT with the dataset, indicating due to the light prob is not collocated with the object, the ground truth lighting is not accurate. I agree on this aspect but insist that due to this issue, the errors in lighting ground truth do not necessarily favor one method or the other. It is simply not reliable for ANY method. In conclusion, I would think it is not justified that the result from this method actually outperforms Neural-PBIR in the presence of the noisy GT lighting, and recommend the authors to rephrase the related language.
[3] On the comparison of results against Neural-PBIR. I would conclude that the proposed method and Neural-PBIR both hit and miss in some aspects over the provided samples, w.r.t. level of details, correctness of specularity and shadows. None is overwhelmingly superior to the other. As a result, the performance of the proposed method is on par with Neural-PBIR over all without significant and meaningful improvement.
Overall, I like the idea and the formulation of the method a lot, but believe there is no significant improvement in the results when compared with baselines. Drawbacks of the method and the paper are also pronounced: long and brittle pipeline (as suggested by Reviewer hgwb), modules (particularly RDM) not sufficiently evaluated, and unjustified claims on the performance. I would refrain from accepting the paper until hearing more from the authors and reviewers.