PoseNet3D: Learning Temporally Consistent 3D Human Pose via Knowledge Distillation

Recovering 3D human pose from 2D joints is a highly unconstrained problem. We\npropose a novel neural network framework, PoseNet3D, that takes 2D joints as\ninput and outputs 3D skeletons and SMPL body model parameters. By casting our\nlearning approach in a student-teacher framework, we avoid using any 3D data\nsuch as paired/unpaired 3D data, motion capture sequences, depth images or\nmulti-view images during training. We first train a teacher network that\noutputs 3D skeletons, using only 2D poses for training. The teacher network\ndistills its knowledge to a student network that predicts 3D pose in SMPL\nrepresentation. Finally, both the teacher and the student networks are jointly\nfine-tuned in an end-to-end manner using temporal, self-consistency and\nadversarial losses, improving the accuracy of each individual network. Results\non Human3.6M dataset for 3D human pose estimation demonstrate that our approach\nreduces the 3D joint prediction error by 18% compared to previous unsupervised\nmethods. Qualitative results on in-the-wild datasets show that the recovered 3D\nposes and meshes are natural, realistic, and flow smoothly over consecutive\nframes.\n

Paper

References (68)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC