Learning What Makes a Difference from Counterfactual Examples and Gradient Supervision

One of the primary challenges limiting the applicability of deep learning is\nits susceptibility to learning spurious correlations rather than the underlying\nmechanisms of the task of interest. The resulting failure to generalise cannot\nbe addressed by simply using more data from the same distribution. We propose\nan auxiliary training objective that improves the generalization capabilities\nof neural networks by leveraging an overlooked supervisory signal found in\nexisting datasets. We use pairs of minimally-different examples with different\nlabels, a.k.a counterfactual or contrasting examples, which provide a signal\nindicative of the underlying causal structure of the task. We show that such\npairs can be identified in a number of existing datasets in computer vision\n(visual question answering, multi-label image classification) and natural\nlanguage processing (sentiment analysis, natural language inference). The new\ntraining objective orients the gradient of a model's decision function with\npairs of counterfactual examples. Models trained with this technique\ndemonstrate improved performance on out-of-distribution test sets.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC