Temporal Consistency Loss for High Resolution Textured and Clothed 3DHuman Reconstruction from Monocular Video
We present a novel method to learn temporally consistent 3D reconstruction of\nclothed people from a monocular video. Recent methods for 3D human\nreconstruction from monocular video using volumetric, implicit or parametric\nhuman shape models, produce per frame reconstructions giving temporally\ninconsistent output and limited performance when applied to video. In this\npaper, we introduce an approach to learn temporally consistent features for\ntextured reconstruction of clothed 3D human sequences from monocular video by\nproposing two advances: a novel temporal consistency loss function; and hybrid\nrepresentation learning for implicit 3D reconstruction from 2D images and\ncoarse 3D geometry. The proposed advances improve the temporal consistency and\naccuracy of both the 3D reconstruction and texture prediction from a monocular\nvideo. Comprehensive comparative performance evaluation on images of people\ndemonstrates that the proposed method significantly outperforms the\nstate-of-the-art learning-based single image 3D human shape estimation\napproaches achieving significant improvement of reconstruction accuracy,\ncompleteness, quality and temporal consistency.\n