Estimating 3D human poses from video is a challenging problem. The lack of 3D\nhuman pose annotations is a major obstacle for supervised training and for\ngeneralization to unseen datasets. In this work, we address this problem by\nproposing a weakly-supervised training scheme that does not require 3D\nannotations or calibrated cameras. The proposed method relies on temporal\ninformation and triangulation. Using 2D poses from multiple views as the input,\nwe first estimate the relative camera orientations and then generate 3D poses\nvia triangulation. The triangulation is only applied to the views with high 2D\nhuman joint confidence. The generated 3D poses are then used to train a\nrecurrent lifting network (RLN) that estimates 3D poses from 2D poses. We\nfurther apply a multi-view re-projection loss to the estimated 3D poses and\nenforce the 3D poses estimated from multi-views to be consistent. Therefore,\nour method relaxes the constraints in practice, only multi-view videos are\nrequired for training, and is thus convenient for in-the-wild settings. At\ninference, RLN merely requires single-view videos. The proposed method\noutperforms previous works on two challenging datasets, Human3.6M and\nMPI-INF-3DHP. Codes and pretrained models will be publicly available.\n
Paper
References (43)
Scroll for more · 31 remaining