HandVoxNet: Deep Voxel-Based Network for 3D Hand Shape and Pose Estimation from a Single Depth Map

3D hand shape and pose estimation from a single depth map is a new and\nchallenging computer vision problem with many applications. The\nstate-of-the-art methods directly regress 3D hand meshes from 2D depth images\nvia 2D convolutional neural networks, which leads to artefacts in the\nestimations due to perspective distortions in the images. In contrast, we\npropose a novel architecture with 3D convolutions trained in a\nweakly-supervised manner. The input to our method is a 3D voxelized depth map,\nand we rely on two hand shape representations. The first one is the 3D\nvoxelized grid of the shape which is accurate but does not preserve the mesh\ntopology and the number of mesh vertices. The second representation is the 3D\nhand surface which is less accurate but does not suffer from the limitations of\nthe first representation. We combine the advantages of these two\nrepresentations by registering the hand surface to the voxelized hand shape. In\nthe extensive experiments, the proposed approach improves over the state of the\nart by 47.8% on the SynHand5M dataset. Moreover, our augmentation policy for\nvoxelized depth maps further enhances the accuracy of 3D hand pose estimation\non real data. Our method produces visually more reasonable and realistic hand\nshapes on NYU and BigHand2.2M datasets compared to the existing approaches.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC