Thank you once again for your valuable feedback on our paper.
1. The time comparison of our method is presented in Table 6 of Appendix A.6 in the original manuscript (for your convenience, we have displayed the table below). Here, we compare the training time per epoch of our method with BP and other state-of-the-art non-BP methods. For the semi-supervised node classification task, all methods are trained for 1000 epochs. During training, we calculate the prediction accuracy on the validation nodes after each epoch and save the model that achieves the highest validation accuracy. This saved model is then used to make predictions on the test nodes. This approach is standard in semi-supervised graph learning. Therefore, the overall training cost is simply one thousand times the numbers shown in Table 6.
2. As noted in our response to W2 and Q3, both BP and the DFA in our method have a complexity of $\mathcal{O}(n^{2})$. Besides, our method introduces an additional complexity of $\mathcal{O}(n^{2})$ due to the pseudo-error generator. As a result, the overall complexity for both our method and BP remains the same $\mathcal{O}(n^{2})$. This explains why our method consistently takes five to ten times longer per epoch compared with BP as shown in Table 6, rather than showing an exponential difference, regardless of dataset size. While BP benefits from decades of research and strong software and hardware support, our method has not yet reached comparable time efficiency of BP. However, our method demonstrates superior time efficiency compared with most non-BP state-of-the-art methods, and the current time cost difference compared with BP is not substantial. Additionally, the parallel update strategy employed by our method for each layer offers considerable potential for parallel computing, which could be further explored to enhance time efficiency.
3. We theoretically derive formulas to integrate the direct feedback alignment mechanism into GNNs, as DFA for fully connected layers is not directly applicable to graph data. Additionally, we provide theoretical proof of the convergence of our proposed method (as shown in Section 4.3 and Appendix A.3). Our work highlights the promising potential of non-BP training in graph deep learning and opens up avenues for tackling some challenges within graphs, which merit further exploration in future research.
If our rebuttal has satisfactorily addressed your concerns, we would greatly appreciate it if you could consider reevaluating the score of our paper. Regardless of your decision, we are truly grateful for your guidance and the time you have invested in reviewing our work.
Thank you again for your attention and support.
Best regards,
Authors
$\newline$
**Table 6: Average running time per epoch (s). For layer-wise training methods like PEPITA, CaFo, FF, and SF, the total time taken by each layer per epoch is reported.**
| Datasets | BP | PEPITA | CaFo+CE | FF+LA | FF+VN | SF | ours |
|----------|---------|---------|---------|--------|--------|---------|---------|
| Cora | 7.56e-3 | 8.73e-3 | 7.61e-1 | 3.14e-1 | 2.83e-1 | 5.49e-2 | 5.66e-2 |
| CiteSeer | 1.06e-2 | 1.11e-2 | 7.68e-1 | 2.59e-1 | 2.61e-1 | 6.88e-2 | 5.68e-2 |
| PubMed | 1.07e-2 | 1.07e-2 | 8.24e-1 | 6.94e-1 | 7.61e-1 | 5.34e-1 | 6.76e-2 |
| Photo | 8.74e-3 | 1.03e-2 | 7.98e-1 | 2.11 | 1.91 | 4.87e-1 | 5.81e-2 |
| Computer | 1.08e-2 | 1.05e-2 | 7.80e-1 | 4.82 | 4.14 | 7.61e-1 | 6.29e-2 |
| Texas | 6.13e-3 | 1.07e-2 | 8.05e-1 | 1.47e-1 | 1.56e-1 | 6.88e-2 | 5.60e-2 |
| Cornell | 5.42e-3 | 1.06e-2 | 7.46e-1 | 1.51e-1 | 1.24e-1 | 3.59e-2 | 5.53e-2 |
| Actor | 9.45e-3 | 1.03e-2 | 7.83e-1 | 6.84e-1 | 6.71e-1 | 2.80e-1 | 5.80e-2 |
| Chameleon| 6.24e-3 | 1.13e-2 | 7.97e-1 | 2.28e-1 | 2.09e-1 | 6.88e-2 | 5.61e-2 |
| Squirrel | 7.77e-3 | 1.20e-2 | 7.78e-1 | 5.82e-1 | 5.05e-1 | 1.21e-1 | 5.79e-2 |