TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation
Remote sensing (RS) vision tasks require extensive labeled data across multiple, interconnected domains. However, current generative data-augmentation frameworks are task-isolated, i.e., each vision task requires training an independent generative model, and ignore the modeling of geographical information and spatial constraints. To address these issues, we propose TerraGen, a unified layout-to-image generation framework that enables flexible, spatially controllable synthesis of RS imagery for various high-level vision tasks, e.g., detection, segmentation, and extraction. In particular, TerraGen introduces a geographic-spatial layout encoder that unifies bounding-box and segmentation mask inputs, combined with a multiscale injection scheme and mask-weighted loss to explicitly encode spatial constraints, from global structures to fine details. Moreover, we construct the first large-scale multi-task RS layout generation dataset and establish a standardized evaluation protocol for this task. Experimental results show that TerraGen achieves the best image-generation quality across diverse tasks. In addition, TerraGen can be used as a universal data-augmentation generator, enhancing downstream task performance significantly and demonstrating robust cross-task generalization in both full-data and few-shot scenarios.