On the Effectiveness of Textual Prompting with Lightweight Fine-Tuning for SAM3 Remote Sensing Segmentation
Remote sensing (RS) segmentation is limited by scarce annotations and domain gaps between overhead and natural imagery, motivating effective adaptation under constrained supervision. SAM3’s concept-driven framework enables mask generation from textual prompts without target-specific modification. We evaluate SAM3 across four RS targets, sources, and resolutions, comparing textual, geometric, and hybrid prompting under lightweight fine-tuning (FT) at increasing supervision levels, and zero-shot (ZS) inference. Results show that combining semantic and geometric cues consistently yields the best performance, while text-only prompting performs worst, particularly for irregular targets, reflecting limited semantic alignment between textual concepts and overhead appearances. Nevertheless, lightly fine-tuned textual prompting offers a favorable performance-effort tradeoff for regular targets. Performance improves sharply from ZS to fine-tuning, followed by diminishing returns, indicating that modest annotation effort may suffice. Persistent Precision–IoU gaps further reveal under-segmentation and boundary errors as dominant failure modes, especially for irregular and less prevalent targets.