**Inconsistencies:**
We thank the reviewer for accurately spotting the naming inconsistency, we have promptly fixed the issue in the manuscript by making $n$ lowercase everywhere. Analogously, we added a figure-level caption to Figure 2 which incorrectly had only two subfigure-level captions. The caption states “*Existing methods accumulate error when cyclically mapping a model through a series of permutations, while $C^2M^3$ correctly maps the model back to the starting point.*”.
Regarding definition 2.1, we used the definition as presented in Git Re-Basin [1]. We believe, however, the reviewer’s comment to be spot on, as the current formulation may hinder clarity. We therefore replaced the $\mathcal{L}(\Theta_A) \approx \mathcal{L}(\Theta_B)$ assumption by instead asking that the two points correspond to weights of neural networks trained to convergence with SGD.
**References**
[1] Ainsworth, Samuel, Jonathan Hayase, and Siddhartha Srinivasa. 2022. “Git Re-Basin: Merging Models modulo Permutation Symmetries.” In *The Eleventh International Conference on Learning Representations (ICLR), 2022*
[2] Jordan, Keller, Hanie Sedghi, Olga Saukh, Rahim Entezari, and Behnam Neyshabur. 2023. “REPAIR: REnormalizing Permuted Activations for Interpolation Repair.” In *The Eleventh International Conference on Learning Representations (ICLR), 2023*
[3] Navon, Aviv, Aviv Shamsian, Ethan Fetaya, Gal Chechik, Nadav Dym, and Haggai Maron. 2023. “Equivariant Deep Weight Space Alignment.” in *The Forty-first International Conference on Machine Learning (ICML), 2024*
[4] Peña, Fidel A. Guerrero, Heitor Rapela Medeiros, Thomas Dubail, Masih Aminbeidokhti, Eric Granger, and Marco Pedersoli. “Re-Basin via Implicit Sinkhorn Differentiation.”, IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR), 2023
[5] Singh, Sidak Pal, and Martin Jaggi. 2020. “Model Fusion via Optimal Transport.” In *Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020 (NeurIPS), 2020*
[6] Horoi, Stefan, Albert Manuel Orozco Camacho, Eugene Belilovsky, and Guy Wolf. 2024. “Harmony in Diversity: Merging Neural Networks with Canonical Correlation Analysis.”, in *The Forty-first International Conference on Machine Learning (ICML), 2024*
[7] Lacoste-Julien, S. (2016). Convergence Rate of Frank-Wolfe for Non-Convex Objectives. *ArXiv, abs/1607.00345*.
[8] Imfeld, Moritz, et al. "Transformer fusion with optimal transport." In *The Twelfth International Conference on Learning Representations (ICLR), 2024*
[9] Verma, Neha, and Maha Elbayad. "Merging text transformer models from different initializations." *arXiv preprint arXiv:2403.00986* (2024).
[10] Marczak, Daniel, et al. "MagMax: Leveraging Model Merging for Seamless Continual Learning." ECCV 2024