Fundamental Coordinate Space for Object 6D Pose Estimation

Estimating the 6D pose of objects, including symmetric ones, is a critical task in computer vision and robotics. Previous correspondence-based methods faced challenges with symmetric objects due to the ambiguities they introduce, necessitating the learning of complex one-to-many correspondences between camera space and object coordinate space. To address this issue, we introduce a novel approach that leverages the concept of fundamental coordinate space. This approach transforms one-to-many correspondences into precise one-to-one correspondences, significantly simplifying the network’s learning process and enhancing its pose estimation performance. Our approach begins by identifying an object’s fundamental coordinate space through a comprehensive pipeline. Subsequently, we develop a coordinate-based attention network to predict dense correspondences between the camera and the fundamental coordinate space. The network employs a fusion module based on attention operations to effectively integrate geometry and texture information at arbitrary query points around the object. Experimental results show that our method surpasses previous state-of-the-art models on both T-LESS and NOCS-REAL datasets, improving the ARMSSD score by 1.4 percentage points on T-LESS and the $5^{\circ }5$ cm score by 6 percentage points on NOCS-REAL, demonstrating its superior performance in 6D pose estimation tasks. Our code is available at https://github.com/wanboyan/FCS.

Paper

Full text

PDF

Fundamental Coordinate Space for Object 6D Pose Estimation

Semantic Scholar · Computer Science · 2024

Abstract

Estimating the 6D pose of objects, including symmetric ones, is a critical task in computer vision and robotics. Previous correspondence-based methods faced challenges with symmetric objects due to the ambiguities they introduce, necessitating the learning of complex one-to-many correspondences between camera space and object coordinate space. To address this issue, we introduce a novel approach that leverages the concept of fundamental coordinate space. This approach transforms one-to-many correspondences into precise one-to-one correspondences, significantly simplifying the network’s learning process and enhancing its pose estimation performance. Our approach begins by identifying an object’s fundamental coordinate space through a comprehensive pipeline. Subsequently, we develop a coordinate-based attention network to predict dense correspondences between the camera and the fundamental coordinate space. The network employs a fusion module based on attention operations to effectively integrate geometry and texture information at arbitrary query points around the object. Experimental results show that our method surpasses previous state-of-the-art models on both T-LESS and NOCS-REAL datasets, improving the ARMSSD score by 1.4 percentage points on T-LESS and the $5^{\circ }5$ cm score by 6 percentage points on NOCS-REAL, demonstrating its superior performance in 6D pose estimation tasks. Our code is available at https://github.com/wanboyan/FCS.

References (50)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC