No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification Problems

In real-world classification tasks, each class often comprises multiple\nfiner-grained "subclasses." As the subclass labels are frequently unavailable,\nmodels trained using only the coarser-grained class labels often exhibit highly\nvariable performance across different subclasses. This phenomenon, known as\nhidden stratification, has important consequences for models deployed in\nsafety-critical applications such as medicine. We propose GEORGE, a method to\nboth measure and mitigate hidden stratification even when subclass labels are\nunknown. We first observe that unlabeled subclasses are often separable in the\nfeature space of deep neural networks, and exploit this fact to estimate\nsubclass labels for the training data via clustering techniques. We then use\nthese approximate subclass labels as a form of noisy supervision in a\ndistributionally robust optimization objective. We theoretically characterize\nthe performance of GEORGE in terms of the worst-case generalization error\nacross any subclass. We empirically validate GEORGE on a mix of real-world and\nbenchmark image classification datasets, and show that our approach boosts\nworst-case subclass accuracy by up to 22 percentage points compared to standard\ntraining techniques, without requiring any prior information about the\nsubclasses.\n

Paper

References (60)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC