Towards Achieving Adversarial Robustness by Enforcing Feature Consistency Across Bit Planes

As humans, we inherently perceive images based on their predominant features,\nand ignore noise embedded within lower bit planes. On the contrary, Deep Neural\nNetworks are known to confidently misclassify images corrupted with\nmeticulously crafted perturbations that are nearly imperceptible to the human\neye. In this work, we attempt to address this problem by training networks to\nform coarse impressions based on the information in higher bit planes, and use\nthe lower bit planes only to refine their prediction. We demonstrate that, by\nimposing consistency on the representations learned across differently\nquantized images, the adversarial robustness of networks improves significantly\nwhen compared to a normally trained model. Present state-of-the-art defenses\nagainst adversarial attacks require the networks to be explicitly trained using\nadversarial samples that are computationally expensive to generate. While such\nmethods that use adversarial training continue to achieve the best results,\nthis work paves the way towards achieving robustness without having to\nexplicitly train on adversarial samples. The proposed approach is therefore\nfaster, and also closer to the natural learning process in humans.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC