Deep Learning in Medical Image Registration: Magic or Mirage?

Classical optimization and learning-based methods are the two reigning paradigms in deformable image registration. While optimization-based methods boast generalizability across modalities and robust performance, learning-based methods promise peak performance, incorporating weak supervision and amortized optimization. However, the exact conditions for either paradigm to perform well over the other are shrouded and not explicitly outlined in the existing literature. In this paper, we make an explicit correspondence between the mutual information of the distribution of per-pixel intensity and labels, and the performance of classical registration methods. This strong correlation hints to the fact that architectural designs in learning-based methods is unlikely to affect this correlation, and therefore, the performance of learning-based methods. This hypothesis is thoroughly validated with state-of-the-art classical and learning-based methods. However, learning-based methods with weak supervision can perform high-fidelity intensity and label registration, which is not possible with classical methods. Next, we show that this high-fidelity feature learning does not translate to invariance to domain shift, and learning-based methods are sensitive to such changes in the data distribution. Finally, we propose a general recipe to choose the best paradigm for a given registration problem, based on these observations.

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer Aeaf5/10 · confidence 3/52024-07-10

Summary

The paper delves into the comparison between classical optimization and learning-based medical image registration methods. Some valuable insights are proposed and the authors propose a general recipe to choose the best paradigm for a given registration problem.

Strengths

1. The motivation behind the work is clear. The observation is in detail. 2. Rich experiments discover the regular and make people rethink the development of learning-based registration methods. 3. The illustrations are clear and easy to understand. 4. Clear conclusions are obtained for registration paradigms in different conditions.

Weaknesses

1. New solutions or deeper insights are missed. 2. Findings in Section 4,5,6 are somewhat valuable, but the relevancy between different phenomena is not very clear. The author failed to point out what should we do or what we may further research for unsupervised or supervised DLIR methods.

Questions

What are the potential directions of future work based on this paper?

Rating

5

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

N/A

Reviewer Xf1e4/10 · confidence 5/52024-07-11

Summary

The work benchmarks the traditional methods and deep learning-based methods for medical image registration and gives a general recipe to choose registration methods.

Strengths

Comprehensive experiments: The authors implement several classical variational and deep learning-based registration models on four public datasets of brain CT/MR images.

Weaknesses

- The training set of deep learning models may not be sufficient because the number of image pairs is obviously not enough, which potentially makes the comparison with classical models unfair. - Medical image registration contains a lot of content, it cannot and should not be simply represented by monomodal registration on four brain image datasets. The authors need to implement more experiments on images of more organs and tissues with monomodal and multimodal registration tasks. - The subtitle “DLIR methods do not generalize across datasets” is almost common sense that networks trained on specific datasets cannot easily generalize to other datasets.

Questions

- Please explicitly clarify the advantages compared to the well-known lean2reg challenge. - Does the conclusion still hold on large datasets? AMOS dataset provides lots of CT and MR images can be used for model training. - Why not submit this work to the dataset and benchmark track?

Rating

4

Confidence

5

Soundness

3

Presentation

3

Contribution

1

Limitations

yes

Reviewer 23gn4/10 · confidence 4/52024-07-12

Summary

This paper discussed about an explicit correspondence between mutual information of the distribution and performance of classification registration methods. The authors argued that this correlation will not be affected by the learning-based methods. They validated this hypothesis on both classical and learning-based registration models and found that the weakly supervised learning-based methods performed high-fidelity intensity and label registration. Besides, they showed that these high-fidelity feature learning methods did not translate to invariance to domain shift, ending with proposing a general approach to select the best paradigm for a registration task.

Strengths

**S1.** The authors experimented on the feature learning approach and their translation to invariancy to domain shift, which is an important aspect in deformable image registration. Also, they covered both classical and learning-based image registration to validate their hypothesis. **S2.** The authors considered out-of-distribution datasets for testing and evaluating performances, which showed promising results regarding generalization.

Weaknesses

**W1. Experimental findings.** 1. I appreciate the authors' finding on getting improved performance for classical methods over learning-based methods. However, I would argue that this is typically not always true considering the reported findings in the learning-based papers, e.g., VoxelMorph (VM), TransMorph (TM), etc. This registration results highly depending on the image pairs that are being considered. For example, image pairs with large age variations won't perform well in the learning-based registration tasks (please check VM, Deepflash, and related papers where they restricted the subject ages to deal with large deformations). Besides, inter-patient registration has a higher probability of getting similar results compared to the atlas-based (or pre-selected template) registration, thoroughly discussed in VM and TM papers. I found this important information missing from their experimental evaluation to support their first hypothesis (Sec. 4). I would encourage the authors to consider experimenting and reporting performances for within-patient, inter-patient, and atlas-to-patient registration results. Also, comparing the anatomical dice scores further justifies the registration models as in the existing SOTA it has been discussed about getting larger variations for critical anatomical structures such as ventricles, pallidum, cerebellum, etc. 2. I found some of the mode's performance very inconsistent with the reported version in the original papers. For example - LKU-Net is achieving more than $~0.925\pm0.025$ dice score where the authors' reported best within subject accuracy is $0.8861\pm0.01$ with an average dice score of $0.7758\pm0.0390$. The same goes for the LapIRN model's performance. There's a large difference between the reported dice in this paper vs the original paper's dice. I understand there might be different biases that might be there in terms of implementation. However, the current reported performance of these models raises questions about the credibility of the adapted implementations, preparation of the pairwise images, hyperparameter setup, etc. I would suggest the authors kindly shed some light on this part. 3. I found the supporting experiments to validate the hypothesis presented in Sec. 6 missing some important information. Did the authors perform an affine transformation on all the datasets and what are the pre-processing steps? Are they considering all 3D volumes for their experiments? **W2. Missing large deformation diffeomorphic and related baselines.** I appreciate the authors for trying to carry out thorough experiments on different classical and learning-based registrations. However, I believe the current hypothesis would be more understandable and justifiable if the authors could initiate some experiments considering large deformation diffeomorphic-based registration methods such as LDDMM [1,2,3], where this kind of method is structured upon time-varying velocity fields, that is proven to be more robust in various deformation-based image analysis tasks. **W3. Overclaimed hypothesis.** Statements in L332-L335 seem overstated. The authors need to discuss and experiment on within-subject/class, inter-subject/class, and atlas/template-based registration to validate the stated hypothesis in that line, which seems to be missing in the current version. Besides, the selection of the DLIR/classical registration method depends on the image analysis tasks that are being performed over image registration. Without verifying the studied registration on different image analysis tasks, it is inappropriate to come to a conclusion stated in L332. **W4. (Minor) Technical writing.** The authors tried to aggregate all their potential findings in a structured way. But reading the whole paper kind of messed me up in understanding what are the actual contribution of this paper compared to the other survey/review papers in this domain other than performing experiments on OOD data. Overall, the presentation is kind of above the borderline but I guess if the authors tried to focus on their storyline and make their findings more clearer that would be great for the readers. For example, I found implementation details in most sections which is kind of redundant. Overall, I appreciate the authors for working on this paper which is very relevant as well as important in the medical imaging domain, specifically in medical image registration. However, the current version of the manuscript lacks some important experimental justification and further experiments. With that being said, the current version of the manuscript is under the threshold of acceptance. However, I am open to reconsidering the initial rating if the above concerns are adequately justified. References ---------------- [1] Yang, Xiao, Roland Kwitt, and Marc Niethammer. "Fast predictive image registration." Deep Learning and Data Labeling for Medical Applications: First International Workshop, LABELS 2016, and Second International Workshop, DLMIA 2016, Held in Conjunction with MICCAI 2016, Athens, Greece, October 21, 2016, Proceedings 1. Springer International Publishing, 2016. [2] Shen, Zhengyang, François-Xavier Vialard, and Marc Niethammer. "Region-specific diffeomorphic metric mapping." Advances in Neural Information Processing Systems 32 (2019). [3] Niethammer, Marc, Roland Kwitt, and Francois-Xavier Vialard. "Metric learning for image registration." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019. [4] Wang, Jian, and Miaomiao Zhang. "Deepflash: An efficient network for learning-based medical image registration." Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020.

Questions

Please see the weaknesses section. I tried to summarize all the findings, concerns, and questions there.

Rating

4

Confidence

4

Soundness

3

Presentation

2

Contribution

2

Limitations

Limitations have been discussed in the paper.

Reviewer tuv26/10 · confidence 3/52024-07-13

Summary

The manuscript investigates the characteristics of two types of registration approach based on traditional variational optimization and deep learning. Experiments revealed a correlation between the mutual information of the distribution of per-pixel intensity and labels, and the performance of classical registration methods. Then the manuscript argues unsupervised deep learning does not improve label matching performance compared to traditional methods, whereas supervised learning methods show improved label matching. Lastly they show learning methods do not generalize well.

Strengths

1. This study reflects original thinking, trying to draw insights on a challenging topic. 2. Experiments design is well motivated to study the hypothesis. 3. The results challenge a claim made by a number of existing studies "learning methods can provide improve label matching when optimized in an unsupervised fashion"

Weaknesses

1. Two of the three claims discussed herein seem trivial. It is not surprised that "Supervised DLIR methods demonstrate enhanced label matching" and "DLIR methods do not generalize across datasets". 2. I feel that the abstract and intro set up a high expectation by saying "we propose a general recipe to choose the best paradigm for a given registration problem, based on these observations." but then again it is a one-sentence recipe in the end that is not surprising to readers "a practitioner should choose DLIR methods only if they have access to a large labeled dataset, and their application is limited to the same dataset distribution. In all other cases, classical optimization-based methods are the more accurate and reliable choice"

Questions

In Fig. 1, shouldn't there also be a correlation within each individual dataset?

Rating

6

Confidence

3

Soundness

4

Presentation

3

Contribution

3

Limitations

While the intro and abstract suggest a study on general registration problems, all experiments are based on brain MRI. It is not clear whether the problem/hypothesis/conclusion is specific to brain MRI.

Authorsrebuttal2024-08-10

Revising rating

We thank you for reading the rebuttal. We would request you to increase your score if you think our responses are satisfactory. If not, we would love an opportunity to discuss your concerns.

Reviewer tuv22024-08-11

Thanks for clarification. I'm still positive about the paper after reading all responses and therefore remain my rating.

Reviewer 23gn2024-08-11

Official Comment by Reviewer 23gn

I thank the reviewers for their response. After reading the rebuttal, I have the following standings — - Some of my concerns regarding LDDMM and the performances of different existing models have been adequately addressed. - The paper's contribution is limited considering the hypothesis that the authors evaluated in this paper. - I agree with my other reviewer that this paper is more aligned with the Dataset and Benchmarking track. Besides, I agree that experiments on datasets other than Brain could reinforce the contribution of their standings. More experiments related to intra-subject, atlas-based registration could further rectify their claims. *After reading all the reviewers' comments and the rebuttal, I think the paper is still on the Borderline (keeping my score as it is), considering the technical contribution, carried out experimentations, overall writing, and the submission track.* I suggest that the authors to address the findings from all reviewers in their revised version.

Reviewer Xf1e2024-08-12

Thanks for the response. The claimed new advantages compared to learn2reg are not convincing to me because these conclusions can also be derived from participants' algorithms. Accordingly, I slightly raise my score.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC