Vidu4D: Single Generated Video to High-Fidelity 4D Reconstruction with Dynamic Gaussian Surfels

Video generative models are receiving particular attention given their ability to generate realistic and imaginative frames. Besides, these models are also observed to exhibit strong 3D consistency, significantly enhancing their potential to act as world simulators. In this work, we present Vidu4D, a novel reconstruction model that excels in accurately reconstructing 4D (i.e., sequential 3D) representations from single generated videos, addressing challenges associated with non-rigidity and frame distortion. This capability is pivotal for creating high-fidelity virtual contents that maintain both spatial and temporal coherence. At the core of Vidu4D is our proposed Dynamic Gaussian Surfels (DGS) technique. DGS optimizes time-varying warping functions to transform Gaussian surfels (surface elements) from a static state to a dynamically warped state. This transformation enables a precise depiction of motion and deformation over time. To preserve the structural integrity of surface-aligned Gaussian surfels, we design the warped-state geometric regularization based on continuous warping fields for estimating normals. Additionally, we learn refinements on rotation and scaling parameters of Gaussian surfels, which greatly alleviates texture flickering during the warping process and enhances the capture of fine-grained appearance details. Vidu4D also contains a novel initialization state that provides a proper start for the warping fields in DGS. Equipping Vidu4D with an existing video generative model, the overall framework demonstrates high-fidelity text-to-4D generation in both appearance and geometry.

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer yVYk5/10 · confidence 4/52024-07-11

Summary

The paper presents Vidu4D, a reconstruction model that can accurately reconstruct 4D (sequential 3D) representations from single generated videos. This method addressing key challenges and enabling high-fidelity virtual content creation. The proposed techniques, such as Dynamic Gaussian Surfels (DGS) and the initialization state, are good contributions that can benefit the field of multi-modal generation and 4D reconstruction.https://openreview.net/

Strengths

1. This paper is well-written. 2. The qualitative results outperform existing methods. 3. The proposed Dynamic Gaussian Surfels (DGS) approach sounds good. It optimizes time-varying warping functions to transform Gaussian surfels from a static to a dynamically warped state, precisely depicting motion and deformation over time.

Weaknesses

- What is your video foundation model? Is it Stable Video Diffusion, SORA, Open SORA, or Your Vidu? If you are using an unreleased foundation video model, is the improvement in qualitative results more due to DGS, or is it caused by the video foundation model? If you utilize Vidu, I believe the author needs to provide the results using open source video foundation models like SVD or Open-SORA. - The quantitative evaluation in the paper is limited to a small set of generated videos. - The performance of Vidu4D on a more diverse and larger dataset of generated videos is not reported, which could limit the generalizability of the findings.

Questions

The paper does not discuss the computational complexity or runtime performance of Vidu4D, which could be an important consideration for practical applications of the method.

Rating

5

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

The paper does not mention how the Vidu4D can be extended or adapted to handle other common challenges in 4D reconstruction, such as occlusions, lighting changes, or complex scene dynamics.

Authorsrebuttal2024-08-13

Thank you again for your time and effort in reviewing our work and providing the constructive comments. Please feel free to let us know if you have any further questions by August 13 AoE, we are more than happy to address them.

Reviewer yVYk2024-08-13

Official Review of Submission3117 by Reviewer yVYk

Thanks for the efforts of the authors for solving some my concerns. I will keep my initial positive rating.

Reviewer 7btN4/10 · confidence 5/52024-07-12

Summary

The paper presents Vidu4D, a reconstruction model that excels in accurately reconstructing 4D (i.e., sequential 3D) representations from single generated videos, addressing challenges associated with non-rigidity and frame distortion. At the core of Vidu4D is a proposed Dynamic Gaussian Surfels (DGS) technique. DGS optimizes time-varying warping functions to transform Gaussian surfels (surface elements) from a static state to a dynamically warped state. This transformation enables a precise depiction of motion and deformation over time. To preserve the structural integrity of surfacealigned Gaussian surfels, the authors design the warped-state geometric regularization based on continuous warping fields for estimating normals. Additionally, the method learns refinements on rotation and scaling parameters of Gaussian surfels, which alleviates texture flickering during the warping process and enhances the capture of fine-grained appearance details.

Strengths

1. The paper shows interesting visual results, although the anonymous link in the submission pdf seems not working correctly. 2. The paper proposes a Banmo-based dynamic 2dgs formulation for 4d reconstruction.

Weaknesses

1. The paper's annotation is cluttered and extremely hard to follow. Sometimes ignoring some symbols in formulation is desirable when too many of them are presented. 2. The real framework section 3.3 is extremely short. The "more details in our appendix" seems to be a false promise? 3. The model hasn't evaluated on realistic scenes, so that it can be compared with other 4dgs methods using their official results. 4. The paper uses banmo like learnable joint representation to drive deformation, however hasn't mentioned the limitations of this kind of methods, i.e., the 4d scene needs to be object centric. 5. No evaluations of latency, it seems the pipeline needs a dynamic neus/nerf, then init 2dgs at the extracted zero level set, which makes the methods very dependable of the stage one and very time consuming.

Questions

Besides above problems, I think the below questions are needed to be addressed: What kind of generated video is used as reference video (better to show the reference mono video). Since for 4d reconstruction, if the reference video doesn't show some parts that are occluded across all frames, it is impossible to reconstruct them. An alternative is to use diffusion prior for novel view supervisions to hallucinate these parts, so what is actually happening here????

Rating

4

Confidence

5

Soundness

3

Presentation

1

Contribution

3

Limitations

I'm confident it will be horrendously challenging for average readers to fully grip the hole picture without a major revision. The paper is written to focus on 4d generation task or lift monocular generated video to 4d, however, 80% of the sections are dedicated to representation. If the author want to present this paper as a sknning-based dynamic 2dgs representation paper, then showing 4d generation only is not enough, 4d reconstruction (w/ many widely used benchmarks) should be used as well. Beside, I hope the author can re-organize the paper no matter it is accepted or not. It is better to put some of the cluttered annotations into appendix, and leave some room for actual pipeline of Vidu4D. The Banmo like joint representation and skinning of 2DGS could be fairly straight forward for people in the field to understand, so no need to put all details in the main paper.

Authorsrebuttal2024-08-13

Thank you again for your time and effort in reviewing our work and providing the constructive comments. Please feel free to let us know if you have any further questions by August 13 AoE, we are more than happy to address them.

Reviewer 1Eme6/10 · confidence 4/52024-07-13

Summary

Video generation models have shown great power recently. Transforming generated videos into 3D/4D representations is important for building a world simulator. This paper proposes an improved 4D reconstruction method from single-generated videos. The key component is the dynamic Gaussian surfels (DGS) technique. Incorporating an initialization stage of a non-rigid warping field, the Vidu4D method produces impressive 4D results with the video generation model Vidu.

Strengths

1. The topic is valuable and interesting to transforming single generated videos into 3D/4D representations. The built 3D/4D representation is more controllable and explicit than a single video. Thus it can be used for rendering more videos with elaborate and customized camera trajectories. Besides, this technique has the potential to be a key component for building a world simulator from video generation models. 2. The provided 4D results show impressive rendering quality, reaching the SOTA performance of this/related field. Besides, the normal looks good, revealing the advantage of modeling geometry from the proposed representation.

Weaknesses

1. The motivation/necessity of building surfels needs to be further strengthened. It is easy to understand building surfels will undoubtedly help reconstruct the surface and geometry. If just considering rendering videos from the built 4D, will a vanilla 4D representation (without improvement on surface reconstruction) be enough? Fig. 4 and Table 1 provide convincing results. However, it is suggested to make it clearer in introducing the motivation, e.g. why better geometry leads to better synthesis. 2. The organization of the method could be improved. For better understanding, it is suggested to first introduce the overall framework of Vidu4D and then demonstrate the dynamic Gaussian Surfels technique. 3. The main method is more like a basic representation/reconstruction approach. Will it still benefit reconstruction from monocular videos captured in real life, not generated videos? Yet, real-life videos have better consistency than generated ones.

Questions

1. The provided results are all object-level 4D ones. Will this method work well on scene-level samples? I know it will be hard to generate the surfels of the background. 2. What about the mesh and depth of the generated 4D representations? 3. One missing related work: Liu I, Su H, Wang X. Dynamic Gaussians Mesh: Consistent Mesh Reconstruction from Monocular Videos[J]. arXiv preprint arXiv:2404.12379, 2024.

Rating

6

Confidence

4

Soundness

3

Presentation

2

Contribution

3

Limitations

The paper has clearly addressed the limitations.

Reviewer crn65/10 · confidence 4/52024-07-13

Summary

The paper proposes a technique called Dynamic Gaussian Surfels to effectively reconstruct 4D reqpresentation from a single generated video. DGS optimizes time-varying warping functions to transform Gaussian surfels and the authors adopt Neural SDF for initialization and proposes a geometry regularization technique to preserve the geometry integrity. Extensive experiments on 30 objects proves the effectiveness of the proposed method.

Strengths

a. The resutls are great, show promising application of the proposed DGS. b. The paper is well writting and easy to follow. c. The authors compare their method with various 4D representations, i.e., skinning and bones, NeRF, Gaussian, and achieve better performance.

Weaknesses

a. The novelty is limited. This work seems to be a combination of LARS (bones and warping), Gaussian Surfels/2DGS (3D representation) and SC-GS/4DGS (refinement). The real novelty might be the geometric regularization and field initialization, although the former is also similar to that used in previous 3D works. b. Lack of evaluation details: The authors evaluated the comparison methods on generated video for novel view synthesis in Table 1, however, there is no gt for generated video's novel view resutls. How did the author conduct the evaluation? c. The experiments is not extensive enough: The authors claim that their methods are designed for generated video, however, I didn't see any special designs, e.g., sovle the potential multi-view inonconsistency in the input video. So, I think it is a general 4D reconstruction method, and it is recommanded to test the proposed methods on commanly used 4D reconstruction datasets (object and scene-level), using standard evaluation metrics.

Questions

See weaknesses. I might adjust the score according to the response from the authors.

Rating

5

Confidence

4

Soundness

2

Presentation

3

Contribution

3

Limitations

Yes

Authorsrebuttal2024-08-12

Further response for Reviewer crn6's feedback

Sincerely thank you for your feedback. We have some further responses w.r.t. (a) 4D reconstruction and (b) special designs for generated videos. **(a) 4D reconstruction** Since our focus is on the 4D reconstruction, to the best of our knowledge, there are no **pose-free** benchmarks for **dynamic scenes**. To address your concern, we build the benchmark by collecting openly available videos from the SORA official webpage. We then strictly compare our method against existing state-of-the-art 4D reconstruction methods. We provide the details and results below. **Benchmark details:** We collect 35 sub-videos from the SORA [1] webpage, including Drone_Ancient_Rome (20 seconds in total, split into 4 sub-videos), Robot_Scene (20 seconds in total, split into 3 sub-videos), Seaside_Aerial_View (20 seconds in total, split into 3 sub-videos), Mountain_Horizontal_View (17 seconds in total, split into 3 sub-videos), Snow_Sakura (17 seconds in total, split into 4 sub-videos), Westworld (25 seconds in total, split into 4 sub-videos), Chrismas_Snowman (17 seconds in total, split into 3 sub-videos), Butterfly_Under_Sea (20 seconds in total, split into 3 sub-videos), Minecraft1 (20 seconds in total, split into 4 sub-videos), Minecraft2 (20 seconds in total, split into 4 sub-videos). **Evaluation method:** We follow the standard pipeline for dynamic reconstruction (Hyper-NeRF, SC-GS, etc), to construct our evaluation setup by selecting every fourth frame as a training frame and designating the middle frame between each pair of training frames as a validation frame. **Results:** | Method | PSNR $\uparrow$ | SSIM $\uparrow$ | LPIPS $\downarrow$ | | :---: | :----: | :---: | :---: | | Deformable-GS | 12.72 | 0.5773 | 0.2861 | | 4D-GS | 12.15| 0.5609 | 0.2926 | | SC-GS | 14.81 | 0.5914 | 0.2420 | | SpacetimeGaussians | 13.24 | 0.5836 | 0.2633 | | Ours without field initialization (Sec. 3.3) | 15.42 | 0.6167 | 0.2268 | | Ours without dual branch refinement (Line 187) | 18.57 | 0.6852 | 0.1945 | | **Ours (full model)** | **19.05** | **0.7323** | **0.1839** | Upon acceptance, we will open-source this benchmark and the corresponding codebase to ensure its reproduction. Besides, since currently there are no pose-free benchmarks for dynamic scenes, we believe this built benchmark is a contribution. Our full model achieves 4.24 PSNR improvement compared to the existing best method. The result proves the superiority of our method and each element we propose (field initialization, dual branch refinement) on the pose-free dynamic scene benchmark. **(b) Special designs for generated videos** As we summarized in the common response, properties of generated videos include both larger-scale aspects (unknown poses and unexpected movement) and small-scale aspects (flickering, floater occlusion). - For larger-scale aspects, we propose the Field initialization stage which provides a proper start for our Dynamic Gaussian Surfels (DGS) regarding both the pose and the movement (please see warping transformation in Eq. 6 of our main paper). Here we'd like to highlight that the field initialization also benefits movement learning since the warping transformation is learned as a continuous field. This design is novel and especially beneficial for generated videos. We will provide more details of the field initialization during revision. - For small-scale aspects, we have proposed the Dual Branch Refinement (Line 187) and provided ablation studies to prove its effectiveness in alleviating flickering. Again we are grateful for your feedback. [1] Video Generation Models as World Simulators.

Reviewer crn62024-08-13

Thanks for the author's response. This work might be a pioneering work of pose-free dynamic scene reconstruction. In this case, I would suggest to design a standard and rigor evaluation protocol using real-world multi-view video captured with camera poses. Currently, the evaluation camera pose is perhaps the same as training camera pose, lacking the changes in views, perhaps limited by generated video. I would take it as a benchmark for 4D reconstruction from generated video instead of a benchmark for general pose-free 4D reconstruction methods. As for the novelty, despite the explanation of the authors, I think this project is without significant technical innovation. But as the first attempt to reconstruct 4D content from generated video, this work is encouraging. Thus, I'll keep my positive score. The authors are suggested to test their work on various video generation models in the future.

Reviewer crn62024-08-10

Feedback

Thanks for the efforts of the authors for solving my concerns. My feedback is as follows: 1. Novelty: a) This paper is not the first work of generating 4D content using text-to-video models. Previously there are many 4D generation works using text-to-video models, such as 4Dfy (CVPR2024), AYG (CVPR2024), Dream-in-4D(CVPR2024), DG4D (arxiv2023), 4DGen(arxiv2023), aniamte124(arxiv2023), etc. Besides, those works generate 360-degree dynamic objects utilizing text-to-video models. In contrast, this work only generates part of the object, which means the invisible part in the input video is missing in generation results, and the novel view synthesis is limited to small camera movement range. b) Based on the above point, I would suggest the authors to claim that they are the first work which deals with 4D reconstruction from generated videos, rather than the first 4D generation work using text-to-video models. c) I reserve my judgment on the novelty of the 4D reconstruction techniques proposed by the author. I don't think the proposed techniques have special design to handle obvious multi-view inconsistency. (In fact, I didn't see the word "multi-view inconsistency" or similar words in the main paper) 3. Please refer to 1 4. The comparison with SC-GS/D-NeRF/ etc. on dataset w/o gt pose has little meaning. They are not specially designed for no gt scenes. It's suggested to compare with works specially designed for pose-free scenes. For comparison on dataset w/ gt pose, considering there are many regularizations integrated in the proposed pipeline, comparable or slightly better performance is expected. So I decide to maintain the score.

Authorsrebuttal2024-08-12

Response for Ethics Review

In our work, we have carefully considered the potential ethical risks associated with generative models, particularly in video content creation. To address these concerns, our model includes robust safety mechanisms designed to screen and prevent any misuse. Furthermore, we have decided to release only the reconstruction code, ensuring that our contribution does not facilitate the generation of content that could raise ethical issues. We believe these measures effectively mitigate potential ethical risks and align with the standards of the community.

Reviewer 1Eme2024-08-13

Thanks for the authors' rebuttal! My concerns have been addressed and I will keep my initial positive rating. The supplemented contents are highly recommended to be added to the final version.

Authorsrebuttal2024-08-13

Response for Ethics Review

We appreciate the reviewer's emphasis on the importance of discussing potential ethical implications. In our work, we have carefully considered the potential risks associated with generative models, particularly in video content creation. Generative models used in video generation pose significant risks, such as the potential for creating deepfakes or other misleading content that could be used for harmful purposes like misinformation, privacy invasion, or defamation. To mitigate these risks, we have chosen to release only the reconstruction code, deliberately avoiding the release of components that could facilitate the generation of content with ethical concerns. This decision ensures that our contribution is focused on advancing reconstruction techniques without enabling the creation of new, potentially harmful video content. We will also add a discussion on broader impacts in the appendix to thoroughly address ethical considerations, such as the risks associated with video generation, and hope this addition will align our work with community standards.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC