Counterfactual GraphLIME-Enabled Explainable Adversarial Defense for Graph-Based Intrusion Detection Using Residual GAN-FGSM Framework
Graph-based Intrusion Detection Systems (IDS) based on Graph Neural Networks (GNNs) such as GCNs and GATs, have demonstrated significant effectiveness in identifying malicious network traffic. However, these models are vulnerable to adversarial attacks, and their lack of transparency complicates the graph-based model classification. To address these challenges, this work proposes a unified framework that combines a novel Residual Generative Adversarial Attack (ResGAN-FGSM) with a counter-factual explanation module for high-order perturbations, alongside Fast Gradient Sign Method (FGSM)-based refinement. This combination generates semantically coherent and highly disruptive adversarial examples, exposing vulnerabilities in GCN and GAT-based IDS systems that prior methods fail to uncover under full white-box access. Additionally, we develop Counterfactual GraphLIME, a novel explanation framework for graph-structured data. This framework iteratively perturbs a node's 1-hop neighborhood until its label flips, identifying the minimal feature changes required to alter predictions, thereby improving transparency and supporting security diagnostics. Through comprehensive experiments on CICIDS2017, we demonstrate that ResGAN-FGSM reduces GNN-based IDS accuracy by up to 30% under modest perturbation budgets <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\epsilon$</tex>, outperforming FGSM alone. The counterfactual explanations reveal sensitive features, such as inter-arrival time and burst rate, providing valuable insights for model hardening.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex