Learning to Regress Bodies from Images using Differentiable Semantic Rendering

Learning to regress 3D human body shape and pose (e.g.~SMPL parameters) from\nmonocular images typically exploits losses on 2D keypoints, silhouettes, and/or\npart-segmentation when 3D training data is not available. Such losses, however,\nare limited because 2D keypoints do not supervise body shape and segmentations\nof people in clothing do not match projected minimally-clothed SMPL shapes. To\nexploit richer image information about clothed people, we introduce\nhigher-level semantic information about clothing to penalize clothed and\nnon-clothed regions of the image differently. To do so, we train a body\nregressor using a novel Differentiable Semantic Rendering - DSR loss. For\nMinimally-Clothed regions, we define the DSR-MC loss, which encourages a tight\nmatch between a rendered SMPL body and the minimally-clothed regions of the\nimage. For clothed regions, we define the DSR-C loss to encourage the rendered\nSMPL body to be inside the clothing mask. To ensure end-to-end differentiable\ntraining, we learn a semantic clothing prior for SMPL vertices from thousands\nof clothed human scans. We perform extensive qualitative and quantitative\nexperiments to evaluate the role of clothing semantics on the accuracy of 3D\nhuman pose and shape estimation. We outperform all previous state-of-the-art\nmethods on 3DPW and Human3.6M and obtain on par results on MPI-INF-3DHP. Code\nand trained models are available for research at https://dsr.is.tue.mpg.de/.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC