Response to Reviewer rVud
Thank you for your helpful suggestions! We are glad that you think the orthogonal attention is "novel" and the two-pathway architecture is innovative. We address your concerns point by point below.
Q1: Empirical analysis of overfitting.
Thanks to your comments about this. We have performed an experiment on the Elasticity benchmark, varying the number of training data, to illustrate the issue of overfitting when data is limited. Our results indicate that our neural operator exhibits a smaller decrease in performance with fewer training data compared to the baseline model Geo-FNO. We have also included another well-acknowledged neural operator, DeepONet, as an additional baseline in our revised version. This addition serves to further highlight the issue and strengthen the motivation behind our work. What's more, we have also experimented with increasing the number of layers to 30 on the Elasticity benchmark. In contrast, Geo-FNO showed a decline in performance when exceeding 4 layers in [1]. This suggests that our proposed neural operator incorporates effective regularization mechanisms, contributing to its superior performance compared to the baseline model.
Q2: Theoretical analysis of orthogonal attention.
Our proposed attention is based on the insight of eigendecomposition. We have also included theoretical support in the form of Eq (17) and Eq (18) in the **appendix A**, which demonstrates that the parametric $\hat{\psi}$ will converge to the top-$k$ principal eigenfunctions of the unknown ground-truth kernel integral operator. We think it's interesting and inspiring to relate our attention to spectral properties and really appreciate your suggestion.
Q3: The impact of the disentangled design.
Thanks for your interest in the disentangled design. We think that the orthogonal attention naturally requires two pathways and their inputs: the eigenfunctions and the input function, making it challenging to replace the disentangled design.
Q4: Comprehensive evaluation in complex settings.
We agree that more complex practical settings can evaluate the neural operator more comprehensively. We are also interested in the performance in practical scenarios. However, we emphasize that our current empirical studies align with the related works in this field and can prove the effectiveness of our method. We will explore more complex tasks in the next version.
We welcome any questions or corrections you may have and sincerely hope for a reconsideration of our paper's score. Your feedback is highly valuable to us.
[1] Alasdair Tran, Alexander Mathews, Lexing Xie, and Cheng Soon Ong. Factorized fourier neural
operators. In The Eleventh International Conference on Learning Representations, 2023