We thank the reviewer for their thoughtful feedback and valuable suggestions.
**W1: Limited Novelty**
We appreciate the reviewer's recognition of the value of our comprehensive benchmarking study. While it is true that our main focus is on systematically evaluating PEs across different GNN architectures, we believe that this type of empirical analysis is highly valuable to the community, especially given the rapid proliferation of PEs in recent years. Our benchmarking framework is designed to fill a gap in the current landscape by providing a unified evaluation across multiple architectures, including MPNNs and GTs.
Regarding SparseGRIT, we acknowledge that it is an extension of GRIT; however, this approach has not been thoroughly explored in prior works, especially within the context of benchmarking PEs, and offers a practical contribution to making GTs more scalable.
**W2: Justification for PE Performance**
We appreciate the reviewer's suggestion for a deeper theoretical analysis of PE performance. In response, we have added a new section titled “Guidelines for Practitioners” to provide more concrete recommendations on selecting PEs based on our findings. Our study focuses on real-world datasets that are commonly used for benchmarking, where it is challenging to directly correlate the theoretical expressiveness of PEs with practical performance. We believe that our empirical results, alongside the new guidelines, can serve as a valuable resource for researchers and practitioners in choosing the most suitable PEs for their specific tasks.
**W3: Marginal Performance Gains**
Our primary goal was not to propose new state-of-the-art models but to provide a thorough benchmarking of PEs across a diverse set of GNN architectures. A key takeaway from our study is that older models leveraging sparse graph topology can achieve performance on par with more complex, recent architectures when the right PE is applied. This is a valuable insight, showing that the choice of PE can be as impactful as architectural advancements. While significant improvements were not the primary focus, our benchmarking results on datasets like PCQM-Contact, where we set new state-of-the-art performance, highlight the potential of optimizing PEs. Notably, using models like Exphormer, we achieved substantial gains across several datasets, demonstrating that well-chosen PEs can unlock performance without the need for overly complex models.
**Q1: Theoretical Justification for Empirical Results**
Our study primarily focuses on empirical evaluations using real-world datasets that are commonly used for benchmarking. It is often challenging to establish a direct connection between the theoretical expressiveness of positional encodings and their practical performance on these datasets. Real-world graphs come with inherent complexities and noise, which can obscure the direct impact of theoretical properties. Nevertheless, we reference recent work that suggests some PEs have higher graph separation power. While these insights can offer some context, our primary aim was to provide empirical guidance for practitioners based on observed results across diverse, real-world scenarios that are used to evaluate new state-of-the-art models.