Harnessing Transferable Adversarial Examples via Multilayer Attention-Guided Spatial Transformations

Transfer-based adversarial attacks are key for evaluating the robustness of deep neural networks (DNNs) in black-box settings, yet their effectiveness is often constrained by limited cross-model transferability. Existing feature-level approaches typically rely on single-layer attention guidance or static perturbation patterns, which restrict adaptability across diverse architectures. In this work, we introduce a unified adversarial framework, named Multi-layer Attention-guided Spatial Transformations (MAT), to exploit class-discriminative cues from multiple feature layers to craft highly transferable adversarial examples. MAT integrates Multi-layer Attention Fusion to capture complementary low-level and high-level semantics from multiple intermediate layers, Attention-guided Augmentation to selectively perturb noncritical regions while preserving semantic integrity, and Spatial Random Transformation to introduce stochastic spatial augmentations to diversify patterns during optimization. Unlike prior methods that use static or layer-specific attention, MAT dynamically adapts feature guidance to the architecture and task, which enhances generalization. We evaluate MAT against eleven state-of-the-art transfer-based attacks across nine CNN-based and Transformer-based architectures on ImageNet. Comprehensive experiments demonstrate that MAT consistently outperforms eleven state-of-the-art transfer-based attacks in both white-box and black-box settings, including against adversarially trained and input preprocessing-based defensive models, while maintaining higher semantic similarity to the original inputs. It highlights the superior adversarial robustness and excellent adaptability of MAT in adversarial machine learning.

Paper

Full text

PDF

Harnessing Transferable Adversarial Examples via Multilayer Attention-Guided Spatial Transformations

Semantic Scholar · Computer Science · 2026

Abstract

Transfer-based adversarial attacks are key for evaluating the robustness of deep neural networks (DNNs) in black-box settings, yet their effectiveness is often constrained by limited cross-model transferability. Existing feature-level approaches typically rely on single-layer attention guidance or static perturbation patterns, which restrict adaptability across diverse architectures. In this work, we introduce a unified adversarial framework, named Multi-layer Attention-guided Spatial Transformations (MAT), to exploit class-discriminative cues from multiple feature layers to craft highly transferable adversarial examples. MAT integrates Multi-layer Attention Fusion to capture complementary low-level and high-level semantics from multiple intermediate layers, Attention-guided Augmentation to selectively perturb noncritical regions while preserving semantic integrity, and Spatial Random Transformation to introduce stochastic spatial augmentations to diversify patterns during optimization. Unlike prior methods that use static or layer-specific attention, MAT dynamically adapts feature guidance to the architecture and task, which enhances generalization. We evaluate MAT against eleven state-of-the-art transfer-based attacks across nine CNN-based and Transformer-based architectures on ImageNet. Comprehensive experiments demonstrate that MAT consistently outperforms eleven state-of-the-art transfer-based attacks in both white-box and black-box settings, including against adversarially trained and input preprocessing-based defensive models, while maintaining higher semantic similarity to the original inputs. It highlights the superior adversarial robustness and excellent adaptability of MAT in adversarial machine learning.

Similar papers

© 2026 NYSGPT2525 LLC