Finding and Removing Clever Hans: Using Explanation Methods to Debug and Improve Deep Models

Contemporary learning models for computer vision are typically trained on\nvery large (benchmark) datasets with millions of samples. These may, however,\ncontain biases, artifacts, or errors that have gone unnoticed and are\nexploitable by the model. In the worst case, the trained model does not learn a\nvalid and generalizable strategy to solve the problem it was trained for, and\nbecomes a 'Clever-Hans' (CH) predictor that bases its decisions on spurious\ncorrelations in the training data, potentially yielding an unrepresentative or\nunfair, and possibly even hazardous predictor. In this paper, we contribute by\nproviding a comprehensive analysis framework based on a scalable statistical\nanalysis of attributions from explanation methods for large data corpora. Based\non a recent technique - Spectral Relevance Analysis - we propose the following\ntechnical contributions and resulting findings: (a) a scalable quantification\nof artifactual and poisoned classes where the machine learning models under\nstudy exhibit CH behavior, (b) several approaches denoted as Class Artifact\nCompensation (ClArC), which are able to effectively and significantly reduce a\nmodel's CH behavior. I.e., we are able to un-Hans models trained on (poisoned)\ndatasets, such as the popular ImageNet data corpus. We demonstrate that ClArC,\ndefined in a simple theoretical framework, may be implemented as part of a\nNeural Network's training or fine-tuning process, or in a post-hoc manner by\ninjecting additional layers, preventing any further propagation of undesired CH\nfeatures, into the network architecture. Using our proposed methods, we provide\nqualitative and quantitative analyses of the biases and artifacts in various\ndatasets. We demonstrate that these insights can give rise to improved, more\nrepresentative and fairer models operating on implicitly cleaned data corpora.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC