Object detection and pose estimation are required for automating the stocking of shelves in retail stores and improving pose estimation accuracy necessitate instance segmentation. However, conventional methods experience difficulties in moving the robot in real time because they have too many parameters, which increases the processing time. In this study, we developed a high-speed instance segmentation method that solves this problem. Specifically, we focused on the fact that the robot's task target is a single object. Consequently, by choosing one detection target in the robot's view (image), we reduce the size required by the deep neural network and accelerate instance segmentation. We used the attention region as our selection method because it does not increase the number of parameters. Further, by using an instance in the attention region, we selectively output only high-precision instances. The results of experiments conducted showed a 12.1 pt improvement in instance segmentation precision and 2.5 times faster execution on a CPU compared with previous methods, as well as a 38.4 pt improvement in pose estimation precision. GRAPHICAL ABSTRACT
Paper
Full text
Selective instance segmentation for pose estimation
Semantic Scholar · Computer Science · 2022
Abstract
Object detection and pose estimation are required for automating the stocking of shelves in retail stores and improving pose estimation accuracy necessitate instance segmentation. However, conventional methods experience difficulties in moving the robot in real time because they have too many parameters, which increases the processing time. In this study, we developed a high-speed instance segmentation method that solves this problem. Specifically, we focused on the fact that the robot's task target is a single object. Consequently, by choosing one detection target in the robot's view (image), we reduce the size required by the deep neural network and accelerate instance segmentation. We used the attention region as our selection method because it does not increase the number of parameters. Further, by using an instance in the attention region, we selectively output only high-precision instances. The results of experiments conducted showed a 12.1 pt improvement in instance segmentation precision and 2.5 times faster execution on a CPU compared with previous methods, as well as a 38.4 pt improvement in pose estimation precision. GRAPHICAL ABSTRACT