Dear Reviewer gca5, We are deeply appreciative of your constructive comment.
Strengths: We are grateful for your encouragement. In future versions (subject to submission size limitations), we will provide a thoroughly segmented dataset to facilitate the reproducibility of our work by all researchers. Additionally, all comparative baseline results can be referenced from FedSMOO, which will further support the reproducibility of our study.
Weakness 1: We appreciate your feedback and recognize that there may have been an issue with our articulation. Our contribution primarily emphasizes being the first to introduce global updates into SAM and dynamic regularization, which, to the best of our knowledge, has not been addressed in previous studies. We certainly do not intend to overlook the significant contributions of FedSMOO, FedLESAM, and FedSpeed regarding the use of dynamic regularization. Instead, we have repeatedly acknowledged these contributions throughout the main text and appendices. For example, see Abstract lines 015-016, Introduction lines 073-075, Appendices A.1 and A.2, as well as the Contributions section. Additionally, we revisit the limitations of dynamic regularization in prior research in Section 2.2. This combination is innovative, and to our knowledge, no study has mentioned this ingenious integration.
Weakness 2: The parameter "\rho" refers to the perturbation learning rate, and "\delta" denotes the global update gradient. We have added more detailed explanations in the main text to clarify these points.
Weakness 3: "theta_i" represents the local model loaded on the client side. We have further refined the notation explanations and added a table of variable definitions in the appendix to enhance clarity.
Weakness 4: The introduction of bounded variance for unit gradients is aimed at facilitating the derivation of Theorem 1, which demonstrates that generalization error and optimization error are geometrically amplified as the local interval expands. This assumption is crucial for supporting the rationale behind our motivation (Section 3). It is also consistent with the assumptions made in related works such as FedSAM and FedLESAM.
[1]. Qu Z, Li X, Duan R, et al. Generalized federated learning via sharpness aware minimization[C]//International conference on machine learning. PMLR, 2022: 18250-18280.
[2]. Fan Z, Hu S, Yao J, et al. Locally Estimated Global Perturbations are Better than Local Perturbations for Federated Sharpness-aware Minimization[C]//Forty-first International Conference on Machine Learning.
Weakness 5: We appreciate your concern regarding the IID testing. In response, we have initiated comprehensive IID testing to address this issue. Specifically, we show that by adjusting the hyperparameters, the performance of FedTOGA can be restored to its original version, ensuring that it will not perform worse than existing algorithms.
Furthermore, we would like to emphasize that our comparison scenarios strictly adhere to the Non-IID experimental settings used in FedSMOO, FedSpeed, and FedLESAM. This alignment with established experimental protocols allows us to compare our results directly with the original experimental data, thereby ensuring the reliability and validity of our findings.
Weakness 6: Thank you for your valuable comments on our paper. "Standing on the shoulders of giants," our algorithm is indeed inspired by FedSMOO, FedCM, FedSpeed, and FedLESAM. We acknowledge that it is challenging to reasonably integrate existing technologies and provide further explanations. However, rethinking and combining established techniques from a completely new perspective is highly significant. Most importantly, our integration significantly reduces the resource requirements of local clients in both FedSMOO and FedLESAM while achieving better results.
In addition, we have introduced the novel concept of Neighborhood perturbation for the first time. This innovative use of gradient caching optimizes the computation of perturbations, a solution that can be effectively applied to existing SAM-based algorithms. For more details, please refer to Appendix B.7. After implementing this technology, there has been a significant improvement in the convergence speed and accuracy of FedSpeed and FedSMOO.
We would also like to kindly remind you that the name of our algorithm is FedTOGA, not FedTOGO.
We appreciate your concerns and hope that our response adequately addresses them and contributes to the further refinement of our manuscript. Should you have any additional questions or require further clarification, please do not hesitate to let us know.