DeepHPS: End-to-end Estimation of 3D Hand Pose and Shape by Learning from Synthetic Depth

Articulated hand pose and shape estimation is an important problem for\nvision-based applications such as augmented reality and animation. In contrast\nto the existing methods which optimize only for joint positions, we propose a\nfully supervised deep network which learns to jointly estimate a full 3D hand\nmesh representation and pose from a single depth image. To this end, a CNN\narchitecture is employed to estimate parametric representations i.e. hand pose,\nbone scales and complex shape parameters. Then, a novel hand pose and shape\nlayer, embedded inside our deep framework, produces 3D joint positions and hand\nmesh. Lack of sufficient training data with varying hand shapes limits the\ngeneralized performance of learning based methods. Also, manually annotating\nreal data is suboptimal. Therefore, we present SynHand5M: a million-scale\nsynthetic dataset with accurate joint annotations, segmentation masks and mesh\nfiles of depth maps. Among model based learning (hybrid) methods, we show\nimproved results on our dataset and two of the public benchmarks i.e. NYU and\nICVL. Also, by employing a joint training strategy with real and synthetic\ndata, we recover 3D hand mesh and pose from real images in 3.7ms.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC