To perform well on unseen and potentially out-of-distribution samples, it is\ndesirable for machine learning models to have a predictable response with\nrespect to transformations affecting the factors of variation of the input.\nHere, we study the relative importance of several types of inductive biases\ntowards such predictable behavior: the choice of data, their augmentations, and\nmodel architectures. Invariance is commonly achieved through hand-engineered\ndata augmentation, but do standard data augmentations address transformations\nthat explain variations in real data? While prior work has focused on synthetic\ndata, we attempt here to characterize the factors of variation in a real\ndataset, ImageNet, and study the invariance of both standard residual networks\nand the recently proposed vision transformer with respect to changes in these\nfactors. We show standard augmentation relies on a precise combination of\ntranslation and scale, with translation recapturing most of the performance\nimprovement -- despite the (approximate) translation invariance built in to\nconvolutional architectures, such as residual networks. In fact, we found that\nscale and translation invariance was similar across residual networks and\nvision transformer models despite their markedly different architectural\ninductive biases. We show the training data itself is the main source of\ninvariance, and that data augmentation only further increases the learned\ninvariances. Notably, the invariances learned during training align with the\nImageNet factors of variation we found. Finally, we find that the main factors\nof variation in ImageNet mostly relate to appearance and are specific to each\nclass.\n
Paper
References (57)
Scroll for more · 38 remaining