Self-Supervised 3D Human Pose Estimation via Part Guided Novel Image Synthesis

Camera captured human pose is an outcome of several sources of variation.\nPerformance of supervised 3D pose estimation approaches comes at the cost of\ndispensing with variations, such as shape and appearance, that may be useful\nfor solving other related tasks. As a result, the learned model not only\ninculcates task-bias but also dataset-bias because of its strong reliance on\nthe annotated samples, which also holds true for weakly-supervised models.\nAcknowledging this, we propose a self-supervised learning framework to\ndisentangle such variations from unlabeled video frames. We leverage the prior\nknowledge on human skeleton and poses in the form of a single part-based 2D\npuppet model, human pose articulation constraints, and a set of unpaired 3D\nposes. Our differentiable formalization, bridging the representation gap\nbetween the 3D pose and spatial part maps, not only facilitates discovery of\ninterpretable pose disentanglement but also allows us to operate on videos with\ndiverse camera movements. Qualitative results on unseen in-the-wild datasets\nestablish our superior generalization across multiple tasks beyond the primary\ntasks of 3D pose estimation and part segmentation. Furthermore, we demonstrate\nstate-of-the-art weakly-supervised 3D pose estimation performance on both\nHuman3.6M and MPI-INF-3DHP datasets.\n

Paper

References (69)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC