We propose to leverage recent advances in reliable 2D pose estimation with\nConvolutional Neural Networks (CNN) to estimate the 3D pose of people from\ndepth images in multi-person Human-Robot Interaction (HRI) scenarios. Our\nmethod is based on the observation that using the depth information to obtain\n3D lifted points from 2D body landmark detections provides a rough estimate of\nthe true 3D human pose, thus requiring only a refinement step. In that line our\ncontributions are threefold. (i) we propose to perform 3D pose estimation from\ndepth images by decoupling 2D pose estimation and 3D pose refinement; (ii) we\npropose a deep-learning approach that regresses the residual pose between the\nlifted 3D pose and the true 3D pose; (iii) we show that despite its simplicity,\nour approach achieves very competitive results both in accuracy and speed on\ntwo public datasets and is therefore appealing for multi-person HRI compared to\nrecent state-of-the-art methods.\n
Paper
References (26)
Scroll for more · 14 remaining