Most existing monocular 3D pose estimation approaches only focus on a single\nbody part, neglecting the fact that the essential nuance of human motion is\nconveyed through a concert of subtle movements of face, hands, and body. In\nthis paper, we present FrankMocap, a fast and accurate whole-body 3D pose\nestimation system that can produce 3D face, hands, and body simultaneously from\nin-the-wild monocular images. The core idea of FrankMocap is its modular\ndesign: We first run 3D pose regression methods for face, hands, and body\nindependently, followed by composing the regression outputs via an integration\nmodule. The separate regression modules allow us to take full advantage of\ntheir state-of-the-art performances without compromising the original accuracy\nand reliability in practice. We develop three different integration modules\nthat trade off between latency and accuracy. All of them are capable of\nproviding simple yet effective solutions to unify the separate outputs into\nseamless whole-body pose estimation results. We quantitatively and\nqualitatively demonstrate that our modularized system outperforms both the\noptimization-based and end-to-end methods of estimating whole-body pose.\n
Paper
References (81)
Scroll for more · 38 remaining