A hybrid classification-regression approach for 3D hand pose estimation using graph convolutional networks

Hand pose estimation is a crucial part of a wide range of augmented reality\nand human-computer interaction applications. Predicting the 3D hand pose from a\nsingle RGB image is challenging due to occlusion and depth ambiguities.\nGCN-based (Graph Convolutional Networks) methods exploit the structural\nrelationship similarity between graphs and hand joints to model kinematic\ndependencies between joints. These techniques use predefined or globally\nlearned joint relationships, which may fail to capture pose-dependent\nconstraints. To address this problem, we propose a two-stage GCN-based\nframework that learns per-pose relationship constraints. Specifically, the\nfirst phase quantizes the 2D/3D space to classify the joints into 2D/3D blocks\nbased on their locality. This spatial dependency information guides this phase\nto estimate reliable 2D and 3D poses. The second stage further improves the 3D\nestimation through a GCN-based module that uses an adaptative nearest neighbor\nalgorithm to determine joint relationships. Extensive experiments show that our\nmulti-stage GCN approach yields an efficient model that produces accurate 2D/3D\nhand poses and outperforms the state-of-the-art on two public datasets.\n

Paper

References (46)

Scroll for more · 34 remaining

Similar papers

© 2026 NYSGPT2525 LLC