Despite being able to capture a range of features of the data, high accuracy\nmodels trained with supervision tend to make similar predictions. This\nseemingly implies that high-performing models share similar biases regardless\nof training methodology, which would limit ensembling benefits and render\nlow-accuracy models as having little practical use. Against this backdrop,\nrecent work has developed quite different training techniques, such as\nlarge-scale contrastive learning, yielding competitively high accuracy on\ngeneralization and robustness benchmarks. This motivates us to revisit the\nassumption that models necessarily learn similar functions. We conduct a\nlarge-scale empirical study of models across hyper-parameters, architectures,\nframeworks, and datasets. We find that model pairs that diverge more in\ntraining methodology display categorically different generalization behavior,\nproducing increasingly uncorrelated errors. We show these models specialize in\nsubdomains of the data, leading to higher ensemble performance: with just 2\nmodels (each with ImageNet accuracy ~76.5%), we can create ensembles with 83.4%\n(+7% boost). Surprisingly, we find that even significantly low-accuracy models\ncan be used to improve high-accuracy models. Finally, we show diverging\ntraining methodology yield representations that capture overlapping (but not\nsupersetting) feature sets which, when combined, lead to increased downstream\nperformance.\n
Paper
References (59)
Scroll for more · 38 remaining