Comments to Authors
Thank you for the detailed response. While the rebuttal was clear and to the point, I remain __skeptical__ about the presented contributions and experimental results of this submission.
### 1. Single Video Demo?
I checked the quality of the video demo, but the authors __only provide the easiest case__ with only a single scene (Chicken in the HyperNeRF dataset). Why haven't the authors shared comprehensive qualitative results? Considering this paper emphasizes the quality of the rendered video, it's essential to closely examine its temporal consistency. Based on the content shared by the authors, it's challenging for me to assess if the paper genuinely substantiates its statements on dynamic neural rendering quality and enhanced temporal coherence.
- __In my own trials with TiNeuVox and rendering a scene from a static viewpoint, the outcome appears more consistent than what's demonstrated in this paper.__ Note that the authors do not compare the results by TiNeuVox in the video demo.
- Additionally, it's baffling why the comparisons in the video demo are __restricted to just one paper, NSFF__. The current approach makes me __suspicious__ of the quality of this submission. This paper is not ready for the publication.
### 2. Lack of experiments to support authors' claims
The experiments, in their current form, don't offer significant insights. Upon revisiting the abstract, the authors mention:
- we propose DynPoint, an algorithm designed to facilitate the __rapid__ synthesis of novel views for unconstrained monocular videos
- our method exhibits strong robustness in effectively handling __long-duration videos__ without learning a canonical representation.
In the authors' position, I would prioritize contrasting our findings with TiNeuVox, which boasts quicker training and rendering in dynamic environments. Even though the authors contend that TiNeuVox doesn't utilize sceneflow (potentially explaining the limited comparative data throughout the paper, supplementary material, and rebuttal), I firmly believe there's a need to set TiNeuVox against all datasets, showcasing full qualitative video results. Without this, __it feels like the authors may have overclaimed the strength of the submission about the rapid rendering.__
Regarding longer video scenarios, reviewer J3AP has similarly noted: _"The paper claim that the method could handle long-duration video, which I agree. However, no experiment is conducted to prove this ability."_ I also checked the additional results that the authors uploaded, but __none of them still proves its strength regarding the lengthy-scenario, in my opinion__.
Here's the key.
- If the authors argue that quantitative metrics are higher than those of previous methods, I would say that it looks trivial to me.
- However, if the authors could provide a video demo similar to DynIBaR presented on their project page, I will absolutely vote for acceptance.
As a reviewer, showing the proper qualitative results is also one important factor to judge the acceptance of the submission.
### 3. Lack of analysis on the usage of consistent depth
It quite makes sense to me that the authors extend the monocular geometry cues for learning the sceneflow. Typically, as the other reviewers also commented, using consistent depth estimation is a meaningful approach to the dynamic NeRF setup. However, it is unclear whether using consistent depth is quite effective. In terms of quantitative results, I found the one. However, in terms of qualitative results, I have no idea. This paper targets _not the static scene but the dynamic scene_. So many ablation studies simply provide the number. I am not that satisfied with such a submission.
_Moreover, if the authors want to claim a clear difference compared to NSFF or the novel contribution, the authors should have provided a specific ablation study to alleviate such concerns._
For example, MonoSDF (Song et al., Neurips 2022) also exploits the monocular geometric cues, such as depth maps and surface normal maps. As you can see in the manuscript, it provides tons of qualitative results with detailed ablation studies that strongly supports the authors claim on the benefit of using the monocular cues.
However, in this paper, even though the quantitative results achieve state-of-the-art performance, I believe that showing qualitative results is much persuasive, typically for the dynamic neural rendering task. I hope that the authors could provide bunch of the qualitative results with clear comparison with various baselines. Otherwise, I would say that this paper is not ready for the publication.