HandFoldingNet: A 3D Hand Pose Estimation Network Using Multiscale-Feature Guided Folding of a 2D Hand Skeleton

With increasing applications of 3D hand pose estimation in various\nhuman-computer interaction applications, convolution neural networks (CNNs)\nbased estimation models have been actively explored. However, the existing\nmodels require complex architectures or redundant computational resources to\ntrade with the acceptable accuracy. To tackle this limitation, this paper\nproposes HandFoldingNet, an accurate and efficient hand pose estimator that\nregresses the hand joint locations from the normalized 3D hand point cloud\ninput. The proposed model utilizes a folding-based decoder that folds a given\n2D hand skeleton into the corresponding joint coordinates. For higher\nestimation accuracy, folding is guided by multi-scale features, which include\nboth global and joint-wise local features. Experimental results show that the\nproposed model outperforms the existing methods on three hand pose benchmark\ndatasets with the lowest model parameter requirement. Code is available at\nhttps://github.com/cwc1260/HandFold.\n

Paper

References (47)

Scroll for more · 35 remaining

Similar papers

© 2026 NYSGPT2525 LLC