Unsupervised Multi-Modal Image Registration via Geometry Preserving Image-to-Image Translation

Many applications, such as autonomous driving, heavily rely on multi-modal\ndata where spatial alignment between the modalities is required. Most\nmulti-modal registration methods struggle computing the spatial correspondence\nbetween the images using prevalent cross-modality similarity measures. In this\nwork, we bypass the difficulties of developing cross-modality similarity\nmeasures, by training an image-to-image translation network on the two input\nmodalities. This learned translation allows training the registration network\nusing simple and reliable mono-modality metrics. We perform multi-modal\nregistration using two networks - a spatial transformation network and a\ntranslation network. We show that by encouraging our translation network to be\ngeometry preserving, we manage to train an accurate spatial transformation\nnetwork. Compared to state-of-the-art multi-modal methods our presented method\nis unsupervised, requiring no pairs of aligned modalities for training, and can\nbe adapted to any pair of modalities. We evaluate our method quantitatively and\nqualitatively on commercial datasets, showing that it performs well on several\nmodalities and achieves accurate alignment.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC