Cars Can't Fly up in the Sky: Improving Urban-Scene Segmentation via Height-driven Attention Networks

This paper exploits the intrinsic features of urban-scene images and proposes\na general add-on module, called height-driven attention networks (HANet), for\nimproving semantic segmentation for urban-scene images. It emphasizes\ninformative features or classes selectively according to the vertical position\nof a pixel. The pixel-wise class distributions are significantly different from\neach other among horizontally segmented sections in the urban-scene images.\nLikewise, urban-scene images have their own distinct characteristics, but most\nsemantic segmentation networks do not reflect such unique attributes in the\narchitecture. The proposed network architecture incorporates the capability\nexploiting the attributes to handle the urban scene dataset effectively. We\nvalidate the consistent performance (mIoU) increase of various semantic\nsegmentation models on two datasets when HANet is adopted. This extensive\nquantitative analysis demonstrates that adding our module to existing models is\neasy and cost-effective. Our method achieves a new state-of-the-art performance\non the Cityscapes benchmark with a large margin among ResNet-101 based\nsegmentation models. Also, we show that the proposed model is coherent with the\nfacts observed in the urban scene by visualizing and interpreting the attention\nmap. Our code and trained models are publicly available at\nhttps://github.com/shachoi/HANet\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC