Energy-based Epistemic Uncertainty for Graph Neural Networks

In domains with interdependent data, such as graphs, quantifying the epistemic uncertainty of a Graph Neural Network (GNN) is challenging as uncertainty can arise at different structural scales. Existing techniques neglect this issue or only distinguish between structure-aware and structure-agnostic uncertainty without combining them into a single measure. We propose GEBM, an energy-based model (EBM) that provides high-quality uncertainty estimates by aggregating energy at different structural levels that naturally arise from graph diffusion. In contrast to logit-based EBMs, we provably induce an integrable density in the data space by regularizing the energy function. We introduce an evidential interpretation of our EBM that significantly improves the predictive robustness of the GNN. Our framework is a simple and effective post hoc method applicable to any pre-trained GNN that is sensitive to various distribution shifts. It consistently achieves the best separation of in-distribution and out-of-distribution data on 6 out of 7 anomaly types while having the best average rank over shifts on \emph{all} datasets.

Paper

Similar papers

Peer review

Reviewer VZ1v6/10 · confidence 3/52024-06-25

Summary

This paper explores the challenges associated with quantifying uncertainty in Graph Neural Networks (GNNs), particularly in domains involving interconnected data such as graphs. The authors propose a novel method called GEBM, which aggregates energy at various structural levels. This approach enhances predictive robustness and improves performance in distinguishing between in-distribution and out-of-distribution data across different anomaly types and datasets.

Strengths

1. The paper introduces a novel method based on energy considerations. The authors provide sufficient background and related work on the topic, making the main idea accessible even to readers unfamiliar with the subject. 2. Theoretical proofs and theorems are adequately presented to support the proposed methodology. The paper includes comprehensive experimental settings and achieves state-of-the-art performance across various tasks. 3. The paper is well-organized and well-written, enhancing clarity and readability.

Weaknesses

1. The experiments predominantly focus on node-level tasks. It would be beneficial to explore how the proposed method performs on graph-level tasks as well. 2. The paper lacks an ablation study specifically on different types of energy used in the GEBM model. Understanding which energy components are critical remains unclear.

Questions

1. From the equations presented, it appears that this method is not limited to graphs. Are there potential challenges in applying this method to tasks beyond graph-related domains?

Rating

6

Confidence

3

Soundness

3

Presentation

3

Contribution

2

Limitations

The limitations are discussed in this paper. GEBM is a post hoc epistemic estimator; it does not improve aleatoric uncertainty or its calibration.

Reviewer Q6TT7/10 · confidence 3/52024-07-08

Summary

The paper defines an integrable (regularized) energy function to capture epistemic uncertainty via energy a pretrained model. The energy function is a function of the logits so the method is a post-hoc model agnostic. The authors define a diffusion-based hierarchical energy propagation (structure agnostic + local diffusion + group diffusion) which both leads to quantification of uncertainty in graphs, and an evidential model prediction.

Strengths

I count the regularization of the energy function and the theory behind it as a strong point of the paper. Also defining an energy-based model for graphs to address uncertainty is a strong starting point for uncertainty quantification on graphs which is really under-explored by the current time. I also see that the authors have provided a complete experimental setup (with some minor exceptions which I addressed in the weaknesses). In total, I find this paper a strong paper, however I believe that it can be stronger with a better flow of the text, and more contribution in theory for the quality of the UQ.

Weaknesses

1. The authors evaluate their method structurally (line 227) via leaving least homophilic nodes or low-page rank centrality as o.o.d. This seems like there is an implicit assumption of homophily in the graph which is not stated anywhere. In other words, I assumed that the propositions, or setup I should see an assumption of "homophily $\ge$ some constant on expectation...". I see that they referred to this assumption in the limitations, but it is better to be mentioned somewhere in other sections as well. 2. *Clearity:* However successful the method is, I do not see a clear intuition on why these three levels of propagation should be combined and why all are aggregated with weight = 1. For this, I expected an intuitive introductory experiment to clearly show what happens to a node before and after each certain diffusion. More importantly, I see theory to show that the regularization $\pm$ diffusion is integrable, or is not infinity anywhere, but I did not find any theory behind why the approach leads to a good uncertainty quantification in the end; like comparing it with an oracle or finding bounds on the distance from unseen ground truth probability. I also did not see any synthetic experiment in this direction which might help a lot. 3. *Experiments:* (1) I see that in scores like ECE and Brier score the method is not the best. Is there any intuitive explanation for why this is the case? I also strongly recommend the authors to mention that in the limitations of the paper. (2) I see the absence of study on models that have diffusion at the probability space; e.g. APPNP. If the GEBM can improve upon these methods then clearly there is some additional information passed in the energy domain. 4. *Minor Typos:* (1) Line 237 the term "fitted" is used twice. More important (2), in Line 163 there should be "for $\boldsymbol{x} \in {Q}_l$" added somewhere to show that the definition of $f(\boldsymbol{x})$ is limited to that polytope.

Questions

1. Building on weaknesses no. 1. (*W1*), do you assume some homophily property like $\mathbb{P}_{v_i \sim v_j}[y_i = y_j] \ge p$? Is heterophily graph a theoretical limitation of your approach? If yes can you elaborate on the theoretical insight behind it? Note that I see the heterophily is mentioned as a limitation but mostly I can not find a sound explanation of why it is other than just leaving it as an assumption. 2. In evaluation with feature perturbations, why didn't the authors use a random XOR shift instead of a total replacement of features? In that case, you have control over the magnitude shift; intuitively I expect the uncertainty to grow with a correlation to the perturbation probability, but here I can just see the endpoints of the experiment I just mentioned -- fully perturbed node and original node. In general, you can also define the feature shift by randomly selecting from the noise or the original features and controlling the randomness as the magnitude of the shift. 3. In Fig. 2. (robust evidential inference), I can not understand why the result is non-trivial. If the graph is homophily, a simple diffusion over predicted probabilities with a strong coefficient can have a significant denoising effect in the prediction. This is especially in case the perturbation is sparse. What is the model evaluated in Fig. 2? Does it have a similar enhancement in robustness for models that already have a diffusion step like APPNP?

Rating

7

Confidence

3

Soundness

3

Presentation

3

Contribution

3

Limitations

I think the limitations are mentioned clearly.

Reviewer nGhJ6/10 · confidence 4/52024-07-13

Summary

This paper introduces a method for post-hoc epistemic uncertainty estimation in logit-based Graph Neural Networks (GNNs) by aggregating energy scores at different levels, including node, local, and group levels. Extensive experiments show the effectiveness of the proposed framework.

Strengths

1. The paper rigorously evaluates the proposed method under various experimental conditions, such as out-of-distribution (OOD) selection, different GNN backbones, and both inductive and transductive evaluation settings. 2. It comprehensively aggregates uncertainties at multiple levels in the graph, including node-level uncertainties, class-specific neighbor information, and propagated energy through diffusion. 3. The manuscript is well-structured, with a clear presentation of concepts, logical flow, and detailed preliminary knowledge.

Weaknesses

1. The paper lacks a detailed discussion on the selection of hyperparameters, especially for the diffusion module $P_A$. Specifics about the parameters $\alpha$ and $t$ mentioned in Appendix C are not sufficiently discussed. Including ablation studies on different graph diffusion architectures, such as label propagation referenced in the Appendix or APPNP used in the GPN paper, would enhance the paper. 2. The paper states that common GNNs suffer from overconfidence due to their similarity to findings on ReLU neural networks[1]. However, literature [2] [3] suggests that predictions from shallow GNNs are typically under-confident. The paper will benefit from evidence on the over-confidence issue of GNNs. 3. Section 4.4 discusses the relationship between energy scores from logit-based classifiers and total evidence in evidential models. The paper lacks an explanation for why the proposed model outperforms evidential models in epistemic uncertainty prediction, particularly how it addresses the feature collapsing issue in density-based models [4]. [1] Matthias Hein, Maksym Andriushchenko, and Julian Bitterwolf. Why relu networks yield high-confidence predictions far away from the training data and how to mitigate the problem. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 41–50, 2019. [2] Wang, Xiao, Hongrui Liu, Chuan Shi, and Cheng Yang. "Be confident! towards trustworthy graph neural networks via confidence calibration." Advances in Neural Information Processing Systems 34 (2021): 23768-23779. [3] Wang, Min, Hao Yang, Jincai Huang, and Qing Cheng. 2024. “Moderate Message Passing Improves Calibration: A Universal Way to Mitigate Confidence Bias in Graph Neural Networks”. Proceedings of the AAAI Conference on Artificial Intelligence 38 (19):21681-89. https://doi.org/10.1609/aaai.v38i19.30167. [4] Mukhoti, Jishnu, Andreas Kirsch, Joost van Amersfoort, Philip HS Torr, and Yarin Gal. "Deep deterministic uncertainty: A simple baseline." arXiv preprint arXiv:2102.11582 (2021).

Questions

1. How does the paper perform inductive training on the GCN backbone when OOD nodes and edges are excluded during the training phase? Does it use graph sampling or data augmentation techniques? 2. In corollary 4.3, what is meant by ' any $x\in \mathbb{R}^d$ ’? Please provide a precise range for $x$ or probability. 3. In the Equation (9), the regularized energy from three structural scales equally contributes to the final energy score. Table 3 shows varying impacts of energy at these scales. Why was the decision made to use equal weighting? 4. There are inconsistencies between some model names mentioned in Section 5.1 and those in the tables. 5. What are the differences in distribution shifts used in this paper compared to those in GPN or GNNSafe, and why did the authors make these changes?

Rating

6

Confidence

4

Soundness

3

Presentation

3

Contribution

2

Limitations

YES

Reviewer VZ1v2024-08-11

Reply to the rebuttal

Thanks for your rebuttal. My concerns have been partially addressed. I would like to maintain the rating.

Authorsrebuttal2024-08-12

We are glad that we could resolve some concerns and want to thank the reviewer again for the time spent on the review, in particular for pointing out interesting directions for future work.

Reviewer nGhJ2024-08-12

Thank you to the authors for their efforts in providing additional experiments and clarifications. They have addressed my concerns, and I have increased my score. Additionally, I agree with most of the points raised in the rebuttal, particularly regarding feature collapsing and under/over-confidence scenarios. I am also interested in exploring the differences and commonalities between energy-based models and evidential-based models.

Authorsrebuttal2024-08-14

We are happy that the reviewer finds the additional ablations and clarifications helpful and we, too, believe it makes the paper stronger. Thank you for the very useful input!

Reviewer Q6TT2024-08-14

Thanks for the detailed reply. My concerns were partially addressed. With the informative reply from the authors, I find this paper an acceptable and strong study. This is why I increase my score.

Authorsrebuttal2024-08-14

We are glad that we could address the reviewer's concerns and believe that our paper benefits from additional experiments prompted by their feedback. Thank you for the very helpful review!

Program Chairsdecision2024-09-25

Decision

Accept (spotlight)

© 2026 NYSGPT2525 LLC