DiagViB-6: A Diagnostic Benchmark Suite for Vision Models in the Presence of Shortcut and Generalization Opportunities

Common deep neural networks (DNNs) for image classification have been shown\nto rely on shortcut opportunities (SO) in the form of predictive and\neasy-to-represent visual factors. This is known as shortcut learning and leads\nto impaired generalization. In this work, we show that common DNNs also suffer\nfrom shortcut learning when predicting only basic visual object factors of\nvariation (FoV) such as shape, color, or texture. We argue that besides\nshortcut opportunities, generalization opportunities (GO) are also an inherent\npart of real-world vision data and arise from partial independence between\npredicted classes and FoVs. We also argue that it is necessary for DNNs to\nexploit GO to overcome shortcut learning. Our core contribution is to introduce\nthe Diagnostic Vision Benchmark suite DiagViB-6, which includes datasets and\nmetrics to study a network's shortcut vulnerability and generalization\ncapability for six independent FoV. In particular, DiagViB-6 allows controlling\nthe type and degree of SO and GO in a dataset. We benchmark a wide range of\npopular vision architectures and show that they can exploit GO only to a\nlimited extent.\n

Paper

References (36)

Scroll for more · 24 remaining

Similar papers

© 2026 NYSGPT2525 LLC