Rebuttal from authors
Thank you for your positive response and we are honored and happy to continue the discussion with you. Regarding your comments above, we have listed and addressed the two concerns you raised below.
## About the comment "there is a theoretically-grounded literature for the convex case seems to contradict your intuition".
In our overall response, we have made it very clear that even for non-convex objectives, the classical federated primal-dual method FedADMM with partial participation have been theoretically proven to converge [X1]. However, despite the theoretical guarantees, our experiments have validated that FedADMM still faces extreme instability and even divergence during partial participation training (our figure 1). Obviously, **pure theoretical analysis can not provide an absolute guarantee** of an algorithm's feasibility. Therefore, it is not appropriate to dismiss the existence of the problem simply because it has been proven in theory. In fact, we are not the first to observe such non-convergence phenomena in experiments; previous classical works [X2,X3,X4] have repeatedly demonstrated the existence of this phenomenon experimentally. [X4] also indicates that the update mismatch between primal and dual variables leads to a "drift", which is related to the ``dual drift" we pointed out in this paper.
[X1] Wang H, Marella S, Anderson J. Fedadmm: A federated primal-dual algorithm allowing partial participation[C]//2022 IEEE 61st Conference on Decision and Control (CDC). IEEE, 2022: 287-294.
[X2] Xu J, Wang S, Wang L, et al. Fedcm: Federated learning with client-level momentum[J]. arXiv preprint arXiv:2106.10874, 2021.
[X3] Baumgart G A, Shin J, Payani A, et al. Not All Federated Learning Algorithms Are Created Equal: A Performance Evaluation Study[J]. arXiv preprint arXiv:2403.17287, 2024.
[X4] Kang H, Kim M, Lee B, et al. FedAND: Federated Learning Exploiting Consensus ADMM by Nulling Drift[J]. IEEE Transactions on Industrial Informatics, 2024.
## About the comment "Your technique should be shown to work in the "simple" convex case".
We cannot agree with this comment. As you mentioned above, nonconvex optimization is more difficult than convex optimization. In fact, **a large number of papers on learning federated primal-dual methods conduct the experiments on non-convex experiments [X1, X2, X3, X4, X5]**.
More importantly, **the "dual drift" issue is discovered in non-convex experiments** by several previous works (mentioned in the answer above), and our technique is aimed at addressing this problem. We have not identified this issue in convex experiments. Therefore, the reviewers have no reason to ask us to implement this work for convex objectives. If the ``dual drift" issue does not indeed exist in convex experiments, then our virtual update technique is not necessary for training convex models. However, since a lot of studies have identified this problem in non-convex experiments, we have validated that our method effectively addresses this issue in non-convex experiments.
[X1] Zhang X, Hong M, Dhople S, et al. Fedpd: A federated learning framework with adaptivity to non-iid data[J]. IEEE Transactions on Signal Processing, 2021, 69: 6055-6070.
[X2] Acar D A E, Zhao Y, Matas R, et al. Federated Learning Based on Dynamic Regularization[C]//International Conference on Learning Representations.
[X3] Sun Y, Shen L, Huang T, et al. FedSpeed: Larger Local Interval, Less Communication Round, and Higher Generalization Accuracy[C]//The Eleventh International Conference on Learning Representations.
[X4] Kang H, Kim M, Lee B, et al. FedAND: Federated Learning Exploiting Consensus ADMM by Nulling Drift[J]. IEEE Transactions on Industrial Informatics, 2024.
[X5] Zhang Y, Tang D. A differential privacy federated learning framework for accelerating convergence[C]//2022 18th International Conference on Computational Intelligence and Security (CIS). IEEE, 2022: 122-126.
We noted that the reviewers show concerns about the naming. One is with the term "dual drift", and the other is with the title. We believe these are very easy to resolve. We would greatly appreciate any suggestions the reviewers might have for the names.