Dissociating spatial frequency reliance from adversarial robustness advantages in neurally guided deep convolutional neural networks

Deep convolutional neural networks (DCNNs) have approached and even surpassed human-level performance on many visual recognition tasks, yet they remain strikingly vulnerable to near-imperceptible perturbations generated by adversarial attacks. Recent research demonstrates that aligning DCNN representations with human visual cortex activity improves adversarial robustness, but the mechanisms driving this advantage are yet to be understood. One hypothesis suggests that neural alignment confers robustness by biasing models away from brittle high-frequency details and towards the low spatial frequencies (LSF). However, recent work indicates that human object recognition critically depends on a narrow, mid-frequency “human channel”. Interestingly, this band was partially preserved in prior LSF-focused studies. In this work, we explicitly investigate whether a spectral bias towards the LSF or the human channel is the primary driver of the adversarial robustness observed in neurally aligned DCNNs. We first show that DCNNs aligned to higher-order regions of the human ventral visual stream systematically increase their reliance on both the LSF and the human channel. However, directly steering DCNNs towards these bands revealed a clear dissociation. Biasing models towards the human channel, either alone or together with the LSF, does not improve robustness and can even impair it. LSF bias produced some robustness gains, but such improvements are modest despite inducing much larger shifts in spatial-frequency reliance than the neurally-aligned DCNNs. Spatial-frequency-biased models overall show little, if any, increase in similarity to human neural representational geometry. Together, our results suggest that altered spatial-frequency reliance is likely an emergent property of learning more human-like representations rather than the primary mechanism by which neural alignment confers adversarial robustness, and motivate the need for future research examining representational properties beyond spatial-frequency biases.

Paper

Similar papers

© 2026 NYSGPT2525 LLC