Neural networks have greatly boosted performance in computer vision by\nlearning powerful representations of input data. The drawback of end-to-end\ntraining for maximal overall performance are black-box models whose hidden\nrepresentations are lacking interpretability: Since distributed coding is\noptimal for latent layers to improve their robustness, attributing meaning to\nparts of a hidden feature vector or to individual neurons is hindered. We\nformulate interpretation as a translation of hidden representations onto\nsemantic concepts that are comprehensible to the user. The mapping between both\ndomains has to be bijective so that semantic modifications in the target domain\ncorrectly alter the original representation. The proposed invertible\ninterpretation network can be transparently applied on top of existing\narchitectures with no need to modify or retrain them. Consequently, we\ntranslate an original representation to an equivalent yet interpretable one and\nbackwards without affecting the expressiveness and performance of the original.\nThe invertible interpretation network disentangles the hidden representation\ninto separate, semantically meaningful concepts. Moreover, we present an\nefficient approach to define semantic concepts by only sketching two images and\nalso an unsupervised strategy. Experimental evaluation demonstrates the wide\napplicability to interpretation of existing classification and image generation\nnetworks as well as to semantically guided image manipulation.\n
Paper
References (56)
Scroll for more · 38 remaining