Learning novel representations of variable sources from multi-modal $\textit{Gaia}$ data via autoencoders
Data Release 3 (DR3) has published for the first time epoch photometry, BP/RP (XP) low-resolution mean spectra, and supervised classification results for millions of variable sources. This extensive dataset offers a unique opportunity to study the variability of these objects by combining multiple data products. In preparation for DR4, we propose and evaluate a machine learning methodology capable of ingesting multiple data products to achieve an unsupervised classification of stellar and quasar variability. A dataset of 4 million DR3 sources was used to train three variational autoencoders (VAEs), which are artificial neural networks (ANNs) designed for data compression and generation. One VAE was trained on XP low-resolution spectra, another on a novel approach based on the distribution of magnitude differences in the band, and the third on folded band light curves. Each source was compressed into 15 numbers, representing the coordinates in a 15-dimensional latent space generated by combining the outputs of these three models. The learned latent representation produced by the ANN effectively distinguishes between the main variability classes present in DR3, as demonstrated through both supervised and unsupervised classification analysis of the latent space. The results highlight a strong synergy between light curves and low-resolution spectral data, emphasising the benefits of combining the different data products. A 2D projection of the latent variables revealed numerous overdensities, most of which strongly correlate with astrophysical properties, showing the potential of this latent space for astrophysical discovery. We show that the properties of our novel latent representation make it highly valuable for variability analysis tasks, including classification, clustering, and outlier detection.
Paper
References (25)
Scroll for more · 13 remaining