Decision-based Black-box Attack Against Vision Transformers via Patch-wise Adversarial Removal

Vision transformers (ViTs) have demonstrated impressive performance and\nstronger adversarial robustness compared to Convolutional Neural Networks\n(CNNs). On the one hand, ViTs' focus on global interaction between individual\npatches reduces the local noise sensitivity of images. On the other hand, the\nneglect of noise sensitivity differences between image regions by existing\ndecision-based attacks further compromises the efficiency of noise compression,\nespecially for ViTs. Therefore, validating the black-box adversarial robustness\nof ViTs when the target model can only be queried still remains a challenging\nproblem. In this paper, we theoretically analyze the limitations of existing\ndecision-based attacks from the perspective of noise sensitivity difference\nbetween regions of the image, and propose a new decision-based black-box attack\nagainst ViTs, termed Patch-wise Adversarial Removal (PAR). PAR divides images\ninto patches through a coarse-to-fine search process and compresses the noise\non each patch separately. PAR records the noise magnitude and noise sensitivity\nof each patch and selects the patch with the highest query value for noise\ncompression. In addition, PAR can be used as a noise initialization method for\nother decision-based attacks to improve the noise compression efficiency on\nboth ViTs and CNNs without introducing additional calculations. Extensive\nexperiments on three datasets demonstrate that PAR achieves a much lower noise\nmagnitude with the same number of queries.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC