One major challenge for monocular 3D human pose estimation in-the-wild is the\nacquisition of training data that contains unconstrained images annotated with\naccurate 3D poses. In this paper, we address this challenge by proposing a\nweakly-supervised approach that does not require 3D annotations and learns to\nestimate 3D poses from unlabeled multi-view data, which can be acquired easily\nin in-the-wild environments. We propose a novel end-to-end learning framework\nthat enables weakly-supervised training using multi-view consistency. Since\nmulti-view consistency is prone to degenerated solutions, we adopt a 2.5D pose\nrepresentation and propose a novel objective function that can only be\nminimized when the predictions of the trained model are consistent and\nplausible across all camera views. We evaluate our proposed approach on two\nlarge scale datasets (Human3.6M and MPII-INF-3DHP) where it achieves\nstate-of-the-art performance among semi-/weakly-supervised methods.\n
Paper
References (57)
Scroll for more · 38 remaining