Deep convolutional networks have been quite successful at various image\nclassification tasks. The current methods to explain the predictions of a\npre-trained model rely on gradient information, often resulting in saliency\nmaps that focus on the foreground object as a whole. However, humans typically\nreason by dissecting an image and pointing out the presence of smaller\nconcepts. The final output is often an aggregation of the presence or absence\nof these smaller concepts. In this work, we propose MACE: a Model Agnostic\nConcept Extractor, which can explain the working of a convolutional network\nthrough smaller concepts. The MACE framework dissects the feature maps\ngenerated by a convolution network for an image to extract concept based\nprototypical explanations. Further, it estimates the relevance of the extracted\nconcepts to the pre-trained model's predictions, a critical aspect required for\nexplaining the individual class predictions, missing in existing approaches. We\nvalidate our framework using VGG16 and ResNet50 CNN architectures, and on\ndatasets like Animals With Attributes 2 (AWA2) and Places365. Our experiments\ndemonstrate that the concepts extracted by the MACE framework increase the\nhuman interpretability of the explanations, and are faithful to the underlying\npre-trained black-box model.\n
Paper
References (29)
Scroll for more · 17 remaining