A Two-Phase Evolutionary Framework to Boost Adversarial Robustness in Neural Networks

Deep neural networks (DNNs) achieve remarkable performance in machine learning, but are vulnerable to adversarial attacks, such as those generated by Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD), which induce misclassifications, endangering safety-critical applications like autonomous driving, medical diagnostics, and cybersecurity. This study proposes a two-phase training framework to enhance adversarial robustness without sacrificing accuracy. The initial phase establishes baseline performance, while the adaptive phase iteratively evolves the dataset by incorporating adversarial synthetic data generated via FGSM and PGD, targeting model vulnerabilities through cycles of training, evaluation, and dataset augmentation. Evaluated on the Modified National Institute of Standards and Technology database (MNIST, $\mathbf{1 0, 0 0 0}$ samples), the Canadian Institute For Advanced Research 10-class dataset (CIFAR-10, $\mathbf{1 0, 0 0 0}$ samples), BrainTumor-7K (7,023 MRIs), and the Phishing Legitimate Dataset (Phishing, $\mathbf{1 0, 0 0 0}$ samples), this framework yields significant robustness gains against FGSM and PGD attacks: 73.73 % on MNIST $(92.71 \pm 0.32 \%$ vs. $18.98 \pm 0.45 \%$ adversarial accuracy), 7.17 % on CIFAR-10 ($41.04 \pm 0.37 \%$ vs. $33.87 \pm 0.41 \%$), 73.22 % on BrainTumor-7K ($89.60 \pm 0.33 \%$ vs. $16.38 \pm 0.42 \%$), and 20.85 % on Phishing ($94.10 \pm 0.22 \%$ vs. $73.25 \pm 0.38 \%$). Furthermore, clean accuracy improves (e.g., 0.20 % MNIST, 3.12 % CIFAR-10), challenging the robustness-accuracy trade-off. This framework provides a scalable and efficient solution for resilient AI, preserving innovative methods for future intellectual property protection in cybersecurity and medical imaging applications

Paper

Full text

PDF

A Two-Phase Evolutionary Framework to Boost Adversarial Robustness in Neural Networks

Semantic Scholar · 2025

Abstract

Deep neural networks (DNNs) achieve remarkable performance in machine learning, but are vulnerable to adversarial attacks, such as those generated by Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD), which induce misclassifications, endangering safety-critical applications like autonomous driving, medical diagnostics, and cybersecurity. This study proposes a two-phase training framework to enhance adversarial robustness without sacrificing accuracy. The initial phase establishes baseline performance, while the adaptive phase iteratively evolves the dataset by incorporating adversarial synthetic data generated via FGSM and PGD, targeting model vulnerabilities through cycles of training, evaluation, and dataset augmentation. Evaluated on the Modified National Institute of Standards and Technology database (MNIST, $\mathbf{1 0, 0 0 0}$ samples), the Canadian Institute For Advanced Research 10-class dataset (CIFAR-10, $\mathbf{1 0, 0 0 0}$ samples), BrainTumor-7K (7,023 MRIs), and the Phishing Legitimate Dataset (Phishing, $\mathbf{1 0, 0 0 0}$ samples), this framework yields significant robustness gains against FGSM and PGD attacks: 73.73 % on MNIST $(92.71 \pm 0.32 %$ vs. $18.98 \pm 0.45 %$ adversarial accuracy), 7.17 % on CIFAR-10 ($41.04 \pm 0.37 %$ vs. $33.87 \pm 0.41 %$), 73.22 % on BrainTumor-7K ($89.60 \pm 0.33 %$ vs. $16.38 \pm 0.42 %$), and 20.85 % on Phishing ($94.10 \pm 0.22 %$ vs. $73.25 \pm 0.38 %$). Furthermore, clean accuracy improves (e.g., 0.20 % MNIST, 3.12 % CIFAR-10), challenging the robustness-accuracy trade-off. This framework provides a scalable and efficient solution for resilient AI, preserving innovative methods for future intellectual property protection in cybersecurity and medical imaging applications

Similar papers

© 2026 NYSGPT2525 LLC