Novel View Synthesis of Dynamic Scenes with Globally Coherent Depths from a Monocular Camera

This paper presents a new method to synthesize an image from arbitrary views\nand times given a collection of images of a dynamic scene. A key challenge for\nthe novel view synthesis arises from dynamic scene reconstruction where\nepipolar geometry does not apply to the local motion of dynamic contents. To\naddress this challenge, we propose to combine the depth from single view (DSV)\nand the depth from multi-view stereo (DMV), where DSV is complete, i.e., a\ndepth is assigned to every pixel, yet view-variant in its scale, while DMV is\nview-invariant yet incomplete. Our insight is that although its scale and\nquality are inconsistent with other views, the depth estimation from a single\nview can be used to reason about the globally coherent geometry of dynamic\ncontents. We cast this problem as learning to correct the scale of DSV, and to\nrefine each depth with locally consistent motions between views to form a\ncoherent depth estimation. We integrate these tasks into a depth fusion network\nin a self-supervised fashion. Given the fused depth maps, we synthesize a\nphotorealistic virtual view in a specific location and time with our deep\nblending network that completes the scene and renders the virtual view. We\nevaluate our method of depth estimation and view synthesis on diverse\nreal-world dynamic scenes and show the outstanding performance over existing\nmethods.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC