Weakly Supervised 3D Human Pose and Shape Reconstruction with Normalizing Flows

Monocular 3D human pose and shape estimation is challenging due to the many\ndegrees of freedom of the human body and thedifficulty to acquire training data\nfor large-scale supervised learning in complex visual scenes. In this paper we\npresent practical semi-supervised and self-supervised models that support\ntraining and good generalization in real-world images and video. Our\nformulation is based on kinematic latent normalizing flow representations and\ndynamics, as well as differentiable, semantic body part alignment loss\nfunctions that support self-supervised learning. In extensive experiments using\n3D motion capture datasets like CMU, Human3.6M, 3DPW, or AMASS, as well as\nimage repositories like COCO, we show that the proposed methods outperform the\nstate of the art, supporting the practical construction of an accurate family\nof models based on large-scale training with diverse and incompletely labeled\nimage and video data.\n

Paper

References (49)

Scroll for more · 37 remaining

Similar papers

© 2026 NYSGPT2525 LLC