Robust and Generalizable Visual Representation Learning via Random Convolutions

While successful for various computer vision tasks, deep neural networks have\nshown to be vulnerable to texture style shifts and small perturbations to which\nhumans are robust. In this work, we show that the robustness of neural networks\ncan be greatly improved through the use of random convolutions as data\naugmentation. Random convolutions are approximately shape-preserving and may\ndistort local textures. Intuitively, randomized convolutions create an infinite\nnumber of new domains with similar global shapes but random local textures.\nTherefore, we explore using outputs of multi-scale random convolutions as new\nimages or mixing them with the original images during training. When applying a\nnetwork trained with our approach to unseen domains, our method consistently\nimproves the performance on domain generalization benchmarks and is scalable to\nImageNet. In particular, in the challenging scenario of generalizing to the\nsketch domain in PACS and to ImageNet-Sketch, our method outperforms\nstate-of-art methods by a large margin. More interestingly, our method can\nbenefit downstream tasks by providing a more robust pretrained visual\nrepresentation.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC