Distinct Views Improve Generalization and Robustness: Combinations of Augmentations With Different Features

Data augmentation is an effective method for improving deep learning model performance. In the vision domain, various augmentation studies have been conducted to enhance generalization ability and robustness against corruption. However, recent augmentation studies have focused on transforming data to be more diverse and challenging. This approach can prevent models from properly learning key features of objects, such as texture and shape. In response, unlike traditional methods that employ a single augmentation strategy, our method simultaneously utilizes three distinct augmentations, each with different characteristics. We transform the images into color-preserving, shape-preserving, and diversity-enhancing views. More specifically, to ensure the model still captures the key factors of visual information, we utilize two feature-preserving views, one with local color (texture) and the other with global shape information. The third view is transformed by an augmentation that enhances diversity. By utilizing these three distinct augmentations, DV (Distinct Views) helps the model effectively learn all the important features of visual information. To further improve robustness against corruption, we incorporate adversarial perturbations into the third (diversity-enhancing) view, unifying additional hardness and diversity. Experimental results show that DV considerably enhances generalization and robustness against corruption, achieving state-of-the-art performance on various image benchmark datasets. Furthermore, we confirm that DV is quite effective even for Transformer-based models, which typically underperform on small datasets.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC