Summary
This paper proposes a general approach for inducing diverse distributional shifts based on graph structure and evaluates the robustness and uncertainty of graph models under these shifts. The authors define several types of distributional shifts based on graph characteristics, such as popularity and locality, and show that these shifts can be quite challenging for existing graph models. They also find that simple models often outperform more sophisticated methods on these challenging shifts. Additionally, the authors explore the trade-offs between the quality of learned representations for the base classification task and the ability to separate nodes under structural distributional shift. Overall, the paper's contributions include a novel approach for creating diverse and challenging distributional shifts for graph datasets, a thorough evaluation of the proposed shifts, and insights into the trade-offs between representation quality and shift detection.
Strengths
Originality: The paper's approach for inducing diverse distributional shifts based on graph structure is novel and fills a gap in the existing literature, which has mainly focused on node features. The authors' proposed shifts are also motivated by real-world scenarios and are synthetically generated, making them a valuable resource for evaluating the robustness and uncertainty of graph models.
Quality: The paper's methodology is rigorous and well-designed, with clear explanations of the proposed shifts and the evaluation metrics used. The authors also provide extensive experimental results that demonstrate the effectiveness of their approach and highlight the challenges that arise when evaluating graph models under distributional shifts.
Clarity: The paper is well-written and easy to follow, with clear explanations of the proposed shifts and the evaluation methodology. The authors also provide helpful visualizations and examples to illustrate their points.
Significance: The paper's contributions are significant and have implications for the development of more robust and reliable decision-making systems based on machine learning. The authors' approach for inducing diverse distributional shifts based on graph structure can be applied to any dataset, making it a valuable resource for researchers and practitioners working on graph learning problems. The insights into the trade-offs between representation quality and shift detection are also important for understanding the limitations of existing graph models and developing more effective ones. Overall, the paper's contributions have the potential to advance the field of graph learning and improve the reliability of machine learning systems.
Weaknesses
Limited scope: The paper focuses solely on node-level problems of graph learning and does not consider other types of graph problems, such as link prediction or graph classification. This limited scope may restrict the generalizability of the proposed approach and its applicability to other types of graph problems.
Synthetic shifts: While the authors' proposed shifts are motivated by real-world scenarios, they are synthetically generated, which may limit their ability to capture the full complexity of real distributional shifts. The paper acknowledges this limitation, but it is still worth noting that the proposed shifts may not fully reflect the challenges that arise in real-world scenarios.
Limited comparison to existing methods: While the paper provides extensive experimental results that demonstrate the effectiveness of the proposed approach, it does not compare the proposed approach to other existing methods for evaluating graph models under distributional shifts. This limits the ability to assess the relative strengths and weaknesses of the proposed approach compared to other approaches.
Questions
Here are some questions and suggestions for the authors:
Can the proposed approach be extended to other types of graph problems, such as link prediction or graph classification? If so, how might the approach need to be modified to accommodate these different types of problems?
How might the proposed approach be adapted to handle real-world distributional shifts, rather than synthetically generated shifts? Are there any limitations to the proposed approach that might make it less effective in handling real-world shifts?
How does the proposed approach compare to other existing methods for evaluating graph models under distributional shifts? Are there any specific strengths or weaknesses of the proposed approach compared to other approaches?
The paper notes that there is a trade-off between the quality of learned representations for the target classification task and the ability to detect distributional shifts using these representations. Can the authors provide more details on this trade-off and how it might impact the effectiveness of the proposed approach in different scenarios?
The paper proposes several different types of distributional shifts based on graph structure. Can the authors provide more details on how these shifts were chosen and whether there are other types of shifts that might be relevant for evaluating graph models?
The paper focuses on evaluating graph models under distributional shifts, but does not provide any guidance on how to modify existing models to improve their robustness to these shifts. Can the authors provide any suggestions or guidelines for modifying existing models to improve their performance under distributional shifts?
The paper notes that the proposed approach can be applied to any dataset, but does not provide any guidance on how to choose the appropriate node property to use as a splitting factor. Can the authors provide any suggestions or guidelines for choosing an appropriate node property for a given dataset?
Overall, the paper presents an interesting and novel approach for evaluating graph models under distributional shifts. However, there are several areas where the authors could provide more details or guidance to help readers better understand and apply the proposed approach.
Rating
7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence
5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.
Limitations
The paper briefly acknowledges some of the limitations of the proposed approach, such as the fact that the synthetic shifts may not fully reflect the complexity of real-world distributional shifts. However, the paper does not provide a detailed discussion of the potential negative societal impact of the work.
Given the technical nature of the paper, it is possible that the authors did not see a direct connection between their work and potential negative societal impacts. However, it is always important for authors to consider the broader implications of their work, especially in fields like machine learning where there is a growing awareness of the potential risks and harms associated with these technologies.
In future work, the authors could consider providing a more detailed discussion of the potential societal impacts of their work, including any ethical or social considerations that may arise from the use of their approach. This could help to ensure that the work is being developed and applied in a responsible and ethical manner.