FedRSClip: Federated Learning for Remote Sensing Scene Classification Using Vision-Language Models

Remote sensing (RS) image classification is essential for various applications, including agricultural monitoring, urban planning, and land use classification. However, RS data are often distributed across multiple institutions, and due to privacy concerns and data-sharing restrictions, leveraging large-scale datasets in a centralized training framework is challenging. Federated learning (FL) offers a promising solution by enabling collaborative model training across distributed data sources without requiring data centralization. Nevertheless, current vision-language models (VLMs), which typically contain billions of parameters, pose significant communication challenges for traditional FL approaches based on model parameter updates as they would incur substantial communication costs. In this article, we propose FedRSCLIP, the first FL framework designed for RS image classification based on a VLM, specifically, Contrastive Language-Image Pre-training (CLIP). FedRSCLIP addresses the challenges of data heterogeneity and large-scale model transmission in federated environments by introducing prompt learning, which optimizes only a small set of tunable parameters. The framework introduces a dual-prompt mechanism (DPM), comprising shared prompts for global knowledge sharing and private prompts for client-specific adaptation. To maintain semantic coherence between shared and private prompts, we propose the dual-prompt alignment constraint (DPAC), which balances global consistency and local adaptability across diverse client distributions. Additionally, to enhance cross-modal representation learning, we introduce the cross-modal feature alignment constraint (CMFAC) to align multimodal features between text and image prompts.

Paper

References (61)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC