In this paper, we address the challenging task of estimating 6D object poses from a single RGB image. Motivated by the deep learning-based object detection methods, we propose a concise and efficient network that integrates 6D object pose parameter estimation into the object detection framework. Furthermore, for more robust estimation to occlusion, a nonlocal self-attention module is introduced. The experimental results show that the proposed method reaches the state-ofthe-art performance on the YCB-video and the Linemod datasets.
Paper
References (42)
Scroll for more · 30 remaining