Due to its all-weather and day-and-night capabilities, synthetic aperture radar imagery is essential for various applications such as disaster management, Earth monitoring, change detection, and target recognition. However, the scarcity of labeled synthetic aperture radar (SAR) data limits the performance of most deep learning algorithms. To address this issue, we propose a self-supervised learning framework based on masked Siamese vision transformers to learn transferable SAR representations, coined SAFE. Our method leverages contrastive learning principles to train a model on unlabeled SAR data, extracting robust features. Rather than introducing a new learning paradigm, the contribution lies in the adaptation of self-supervised learning to SAR data through modality-aware design choices. SAFE is applicable across multiple SAR acquisition modes and resolutions. We introduce tailored data augmentation techniques specific to SAR imagery, such as subaperture decomposition and despeckling, designed to preserve SAR-consistent invariances. Comprehensive evaluations on various downstream tasks, including few-shot classification, segmentation, visualization, and pattern detection, demonstrate the effectiveness and versatility of the proposed approach. In particular, cross-dataset evaluations show that the learned representations transfer across different sensors and acquisition conditions. Our network yields strong performance in few-shot classification and segmentation tasks, even without being trained on the sensors used for evaluation, highlighting its potential for transferable representation learning in SAR applications.