Learning Positive-Incentive Point Sampling in Neural Implicit Fields for Object Pose Estimation

Learning neural implicit fields of 3D shapes is a rapidly emerging field that enables shape representation at arbitrary resolutions. Due to the flexibility, neural implicit fields have succeeded in many research areas, including shape reconstruction, novel view image synthesis, and more recently, object pose estimation. Neural implicit fields enable learning dense correspondences between the camera space and the object’s canonical space – including unobserved regions in camera space – significantly boosting object pose estimation performance in challenging scenarios like highly occluded objects and novel shapes. Despite progress, predicting canonical coordinates for unobserved camera-space regions remains challenging due to the lack of direct observational signals. This necessitates heavy reliance on the model’s generalization ability, resulting in high uncertainty. Consequently, densely sampling points across the entire camera space may yield inaccurate estimations that hinder the learning process and compromise performance. To alleviate this problem, we propose a method combining an SO(3)-equivariant convolutional implicit network and a positive-incentive point sampling (<monospace>PIPS</monospace>) strategy. The SO(3)-equivariant convolutional implicit network estimates point-level attributes with SO(3)-equivariance at arbitrary query locations, demonstrating superior performance compared to most existing baselines. The <monospace>PIPS</monospace> strategy dynamically determines sampling locations based on the input, thereby boosting the network’s accuracy and training efficiency. The <monospace>PIPS</monospace> strategy is implemented with a <monospace>PIPS</monospace> estimation network which generates sparse sample points with distinctive features capable of determining all object pose DoFs with high certainty. To collect the training data of the <monospace>PIPS</monospace> estimation network, we propose to automatically generate the pseudo ground-truth with a teacher model. Our method outperforms the state-of-the-art on three pose estimation datasets. It achieves 0.63 in the <inline-formula><tex-math notation="LaTeX">$5^{\circ }2$</tex-math><alternatives><mml:math><mml:mrow><mml:msup><mml:mn>5</mml:mn><mml:mo>∘</mml:mo></mml:msup><mml:mn>2</mml:mn></mml:mrow></mml:math><inline-graphic xlink:href="shi-ieq1-3647829.gif"/></alternatives></inline-formula> cm metric on NOCS-REAL275, 0.62 in the <inline-formula><tex-math notation="LaTeX">$5^{\circ }5$</tex-math><alternatives><mml:math><mml:mrow><mml:msup><mml:mn>5</mml:mn><mml:mo>∘</mml:mo></mml:msup><mml:mn>5</mml:mn></mml:mrow></mml:math><inline-graphic xlink:href="shi-ieq2-3647829.gif"/></alternatives></inline-formula> cm metric on ShapeNet-C, and 77.3 in the AR metric on LineMOD-O. Notably, it demonstrates significant improvements in challenging scenarios, such as objects captured with unseen pose, high occlusion, novel geometry, and severe noise.

Paper

Similar papers

© 2026 NYSGPT2525 LLC