Sparse Steerable Convolutions: An Efficient Learning of SE(3)-Equivariant Features for Estimation and Tracking of Object Poses in 3D Space
As a basic component of SE(3)-equivariant deep feature learning, steerable\nconvolution has recently demonstrated its advantages for 3D semantic analysis.\nThe advantages are, however, brought by expensive computations on dense,\nvolumetric data, which prevent its practical use for efficient processing of 3D\ndata that are inherently sparse. In this paper, we propose a novel design of\nSparse Steerable Convolution (SS-Conv) to address the shortcoming; SS-Conv\ngreatly accelerates steerable convolution with sparse tensors, while strictly\npreserving the property of SE(3)-equivariance. Based on SS-Conv, we propose a\ngeneral pipeline for precise estimation of object poses, wherein a key design\nis a Feature-Steering module that takes the full advantage of\nSE(3)-equivariance and is able to conduct an efficient pose refinement. To\nverify our designs, we conduct thorough experiments on three tasks of 3D object\nsemantic analysis, including instance-level 6D pose estimation, category-level\n6D pose and size estimation, and category-level 6D pose tracking. Our proposed\npipeline based on SS-Conv outperforms existing methods on almost all the\nmetrics evaluated by the three tasks. Ablation studies also show the\nsuperiority of our SS-Conv over alternative convolutions in terms of both\naccuracy and efficiency. Our code is released publicly at\nhttps://github.com/Gorilla-Lab-SCUT/SS-Conv.\n
Paper
References (27)
Scroll for more · 15 remaining