Weaknesses
**W1. Experimental findings.**
1. I appreciate the authors' finding on getting improved performance for classical methods over learning-based methods. However, I would argue that this is typically not always true considering the reported findings in the learning-based papers, e.g., VoxelMorph (VM), TransMorph (TM), etc. This registration results highly depending on the image pairs that are being considered. For example, image pairs with large age variations won't perform well in the learning-based registration tasks (please check VM, Deepflash, and related papers where they restricted the subject ages to deal with large deformations). Besides, inter-patient registration has a higher probability of getting similar results compared to the atlas-based (or pre-selected template) registration, thoroughly discussed in VM and TM papers. I found this important information missing from their experimental evaluation to support their first hypothesis (Sec. 4). I would encourage the authors to consider experimenting and reporting performances for within-patient, inter-patient, and atlas-to-patient registration results. Also, comparing the anatomical dice scores further justifies the registration models as in the existing SOTA it has been discussed about getting larger variations for critical anatomical structures such as ventricles, pallidum, cerebellum, etc.
2. I found some of the mode's performance very inconsistent with the reported version in the original papers. For example - LKU-Net is achieving more than $~0.925\pm0.025$ dice score where the authors' reported best within subject accuracy is $0.8861\pm0.01$ with an average dice score of $0.7758\pm0.0390$. The same goes for the LapIRN model's performance. There's a large difference between the reported dice in this paper vs the original paper's dice. I understand there might be different biases that might be there in terms of implementation. However, the current reported performance of these models raises questions about the credibility of the adapted implementations, preparation of the pairwise images, hyperparameter setup, etc. I would suggest the authors kindly shed some light on this part.
3. I found the supporting experiments to validate the hypothesis presented in Sec. 6 missing some important information. Did the authors perform an affine transformation on all the datasets and what are the pre-processing steps? Are they considering all 3D volumes for their experiments?
**W2. Missing large deformation diffeomorphic and related baselines.** I appreciate the authors for trying to carry out thorough experiments on different classical and learning-based registrations. However, I believe the current hypothesis would be more understandable and justifiable if the authors could initiate some experiments considering large deformation diffeomorphic-based registration methods such as LDDMM [1,2,3], where this kind of method is structured upon time-varying velocity fields, that is proven to be more robust in various deformation-based image analysis tasks.
**W3. Overclaimed hypothesis.** Statements in L332-L335 seem overstated. The authors need to discuss and experiment on within-subject/class, inter-subject/class, and atlas/template-based registration to validate the stated hypothesis in that line, which seems to be missing in the current version. Besides, the selection of the DLIR/classical registration method depends on the image analysis tasks that are being performed over image registration. Without verifying the studied registration on different image analysis tasks, it is inappropriate to come to a conclusion stated in L332.
**W4. (Minor) Technical writing.** The authors tried to aggregate all their potential findings in a structured way. But reading the whole paper kind of messed me up in understanding what are the actual contribution of this paper compared to the other survey/review papers in this domain other than performing experiments on OOD data. Overall, the presentation is kind of above the borderline but I guess if the authors tried to focus on their storyline and make their findings more clearer that would be great for the readers. For example, I found implementation details in most sections which is kind of redundant.
Overall, I appreciate the authors for working on this paper which is very relevant as well as important in the medical imaging domain, specifically in medical image registration. However, the current version of the manuscript lacks some important experimental justification and further experiments. With that being said, the current version of the manuscript is under the threshold of acceptance. However, I am open to reconsidering the initial rating if the above concerns are adequately justified.
References
----------------
[1] Yang, Xiao, Roland Kwitt, and Marc Niethammer. "Fast predictive image registration." Deep Learning and Data Labeling for Medical Applications: First International Workshop, LABELS 2016, and Second International Workshop, DLMIA 2016, Held in Conjunction with MICCAI 2016, Athens, Greece, October 21, 2016, Proceedings 1. Springer International Publishing, 2016.
[2] Shen, Zhengyang, François-Xavier Vialard, and Marc Niethammer. "Region-specific diffeomorphic metric mapping." Advances in Neural Information Processing Systems 32 (2019).
[3] Niethammer, Marc, Roland Kwitt, and Francois-Xavier Vialard. "Metric learning for image registration." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019.
[4] Wang, Jian, and Miaomiao Zhang. "Deepflash: An efficient network for learning-based medical image registration." Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020.