Thank you for the review comments!
> Convergence analysis should be conducted to better justify the proposed method, e.g. the norm of A and B, like in figure 1.
Thank you for the suggestion. We have added visualization of update magnitudes of the proposed method in Appendix A.7 (Figure 3 and 4). Further, to make it more clear to visualize the difference between different methods, we plot the relative update magnitude (update norm divided by parameter norm) which is a better indicator of how much the parameters are updated. The figures show that for conventional optimizers, factor A barely changes, while LoRA-RITE is able to learn factor A effectively.
Additionally, we have added the training loss curves in Appendix A.6. The figures show that the training loss of LoRA-RITE decreases faster than the baseline methods.
> The analysis over experimental results is limited, e.g., for some datasets, the proposed method demonstrates significant performance gain compared to LoRA (Adam), what is the property of the dataset such that the proposed optimization can result in such improvement?
We are still continually investigating the property of the dataset where LoRA-RITE can have significant performance gain. One observation we have made so far is that LoRA-RITE usually performs better when there is a gap between the baseline methods and full finetuning, while the improvement is usually milder if their performance is already close. This can be seen from the tables below.
| Gemma-2B | CauseEffectClassification | CoreferenceResolution | Title Generation | DatatoText | Global |
| --------------- | ------------------------- | --------------------- | ---------------- | ---------- | ----------- |
| Adam | 58.93 | 77.06 | 51.30 | 55.52 | 50.51/74.54 |
| ScaledAdam | 58.71 | 77.55 | 51.16 | 55.69 | 49.40/74.01 |
| Lamb | 60.97 | 80.69 | 52.26 | 55.85 | 53.53/76.43 |
| LoRA-RITE | 61.26 | 82.02 | 52.26 | 55.98 | 55.11/77.12 |
| Full Finetuning | 62.77 | 83.50 | 52.58 | 56.09 | 60.22/79.19 |
In the table above, we observe a larger gap between full finetuning and lora in CausalEffectClassification, CorefereferenceResolution and Global, and the gap between LoRA-RITE and baselines is also larger on those tasks. Similar observations can be made in the following table, except the second and fourth column where LoRA-RITE outperforms Full Finetuning, which is probably due to some regularization effect.
| Gemma-2B | Hellaswag | ArcChallenge | gsm8k | openbookqa |
| --------------- | --------- | ------------ | ----- | ---------- |
| Adam | 83.76 | 45.31 | 24.26 | 64.0 |
| ScaledAdam | 83.52 | 45.22 | 23.96 | 64.8 |
| Lamb | 86.60 | 47.35 | 26.76 | 68.0 |
| LoRA-RITE | 87.28 | 49.06 | 30.10 | 68.8 |
| Full Finetuning | 87.33 | 43.14 | 32.6 | 67.0 |
> As analyzed in at line 322, from time complexity perspective, the proposed method is r times slower than Adam, but in Table 4, it is only 8% slower, it will be interesting to elaborate more on the inconsistency.
The training speed measured in Table 4 includes the time for the model’s forward and backward passes. Since the complexity of the forward and backward passes is at least $\Omega(nm)$, the overhead of LoRA-RITE is relatively small. We have added these explanations to the new version of the paper.