High-quality element distribution maps enable pre-cise analysis of Old Master paintings. These maps are typically produced by Macro X-ray Fluorescence (MA-XRF) scanning, a non-invasive technique for elemental imaging of flat surfaces. However, MA-XRF faces a trade-off between resolution and acquisition time, making high-resolution (HR) scans impractical for large artworks. Super-resolution MA-XRF mitigates this by enhancing scan quality while reducing acquisition time. This paper introduces a deep learning framework for MA-XRF super-resolution that removes the need for paired HR MA-XRF training data by leveraging RGB images to model cross-modal dependencies. Our approach is specifically tailored for MA-XRF, an important feature as RGB and MA-XRF data lack a common spectral domain. We introduce self-supervised adversarial training, where the discriminator learns from patches across modalities, guiding the generator toward realistic MA-XRF reconstructions. Additionally, our method enforces physical consistency via network design and enhances training through pseudo-real data augmentation. Experiments on Old Master paintings show our method outperforms state-of-the-art MA-XRF super-resolution techniques, demonstrating the need for tailored solutions as existing approaches from other domains do not generalize effectively to this task.