Non-uniform Image Down-sampling for Efficient Dense Perception

Dense perception tasks such as Semantic Segmentation, Depth Estimation and Object Detection can provide rich perceptual information for self-driving vehicles. As the Deep Neural Network is becoming a trend for these tasks, most of these networks face the great challenge of taking a balance between performance and speed. Typically, downsampling high-resolution input images before feeding it into the network is an effective way to reduce computation overhead but it also sacrifices accuracy. Based on the fact that information is not uniformly distributed in the input image, it might be helpful to downsample the image using a spatial Non-Uniform Sampling (NUS) method, so as to minimize the information loss in the downsampling process. In this paper, we proposed a general and end-to-end trainable NUS-based framework coined DiffeoSTN for dense prediction tasks. Importantly, we reveal that the key ingredient to obtain superior performance through end-to-end training is a reverse (up)sampler corresponding to the NUS downsampler. Specifically, we leverage an intuitive force attraction downsampler and a fixed-point upsampler in the framework. We demonstrate consistently favorable performance compared to non-end-to-end-trained and uniform samplers on multiple datasets.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC