SALInst: Spatial Affinity Learning for Remote Sensing Instance Segmentation With Box Supervision

Remote sensing instance segmentation plays a vital role in geographic information systems, and it is crucial for Internet of Things applications such as smart city infrastructure management, traffic monitoring, and autonomous driving systems. Although fully supervised methods have achieved promising accuracy, they rely heavily on large amounts of pixel-level annotations, which are extremely costly to obtain for high-resolution remote sensing images. Box-supervised instance segmentation typically leverages horizontal bounding boxes as weak supervision signals, significantly reducing the annotation burden. However, mask prediction under box supervision faces challenges due to limited utilization of spatial information, including the lack of geometric details, the disconnect between spatial localization and pixel-level prediction, and the neglect of spatial priors in traditional pairwise affinity loss. To address these issues, this article proposes SALInst, a spatial affinity learning framework for box-supervised remote sensing instance segmentation. Specifically, a spatial information enhancement module and a dual-stream residual gate fusion mechanism are designed to strengthen spatial awareness and semantic coherence of mask features. Furthermore, by leveraging the spatial constraint prior, we propose a spatial affinity loss based on Gaussian kernel and total variation loss to reinforce the spatial consistency of predicted masks. Extensive experiments on the iSAID and NWPU VHR-10 datasets demonstrate that SALInst outperforms existing box-supervised methods while narrowing the performance gap between weakly and fully supervised instance segmentation.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC