A precise understanding of why units in an artificial network respond to\ncertain stimuli would constitute a big step towards explainable artificial\nintelligence. One widely used approach towards this goal is to visualize unit\nresponses via activation maximization. These synthetic feature visualizations\nare purported to provide humans with precise information about the image\nfeatures that cause a unit to be activated - an advantage over other\nalternatives like strongly activating natural dataset samples. If humans indeed\ngain causal insight from visualizations, this should enable them to predict the\neffect of an intervention, such as how occluding a certain patch of the image\n(say, a dog's head) changes a unit's activation. Here, we test this hypothesis\nby asking humans to decide which of two square occlusions causes a larger\nchange to a unit's activation. Both a large-scale crowdsourced experiment and\nmeasurements with experts show that on average the extremely activating feature\nvisualizations by Olah et al. (2017) indeed help humans on this task ($68 \\pm\n4$% accuracy; baseline performance without any visualizations is $60 \\pm 3$%).\nHowever, they do not provide any substantial advantage over other\nvisualizations (such as e.g. dataset samples), which yield similar performance\n($66\\pm3$% to $67 \\pm3$% accuracy). Taken together, we propose an objective\npsychophysical task to quantify the benefit of unit-level interpretability\nmethods for humans, and find no evidence that a widely-used feature\nvisualization method provides humans with better "causal understanding" of unit\nactivations than simple alternative visualizations.\n