CLRCNet: Cascaded Low-rank Convolutions for Semantic Segmentation in Real-time

Recent work has shown that deep convolutional neural networks have made immensely successful in many computer vision tasks such as semantic image segmentation, which can be thought as a complexed localization and classification problem. However, due to the limitation of computing cost and memory, most existing models are difficult to deploy on mobile devices. It is also an arduous task to get more semantic information from the feature map of downsampling. In this paper, we introduce cascaded low-rank convolutions network (CLRCNet) which is an efficient neural network by using multiple low-rank layers. The cascaded low-rank layers are used to reduce computational complexity. A pooling operation in network units is introduced to learn more contextual information during training. A large number of experiments show that the method has better performance than other network structures. Our network struct attains mean intersection over union (mIOU) of 63.3% on Cityscapes dataset at 76.9 frames per second on $512 \times 1024$ resolution.

Paper

Full text

PDF

CLRCNet: Cascaded Low-rank Convolutions for Semantic Segmentation in Real-time

Semantic Scholar · Computer Science · 2019

Abstract

Recent work has shown that deep convolutional neural networks have made immensely successful in many computer vision tasks such as semantic image segmentation, which can be thought as a complexed localization and classification problem. However, due to the limitation of computing cost and memory, most existing models are difficult to deploy on mobile devices. It is also an arduous task to get more semantic information from the feature map of downsampling. In this paper, we introduce cascaded low-rank convolutions network (CLRCNet) which is an efficient neural network by using multiple low-rank layers. The cascaded low-rank layers are used to reduce computational complexity. A pooling operation in network units is introduced to learn more contextual information during training. A large number of experiments show that the method has better performance than other network structures. Our network struct attains mean intersection over union (mIOU) of 63.3% on Cityscapes dataset at 76.9 frames per second on $512 \times 1024$ resolution.

Similar papers

© 2026 NYSGPT2525 LLC