For machine learning models to be most useful in numerous sociotechnical\nsystems, many have argued that they must be human-interpretable. However,\ndespite increasing interest in interpretability, there remains no firm\nconsensus on how to measure it. This is especially true in representation\nlearning, where interpretability research has focused on "disentanglement"\nmeasures only applicable to synthetic datasets and not grounded in human\nfactors. We introduce a task to quantify the human-interpretability of\ngenerative model representations, where users interactively modify\nrepresentations to reconstruct target instances. On synthetic datasets, we find\nperformance on this task much more reliably differentiates entangled and\ndisentangled models than baseline approaches. On a real dataset, we find it\ndifferentiates between representation learning methods widely believed but\nnever shown to produce more or less interpretable models. In both cases, we ran\nsmall-scale think-aloud studies and large-scale experiments on Amazon\nMechanical Turk to confirm that our qualitative and quantitative results\nagreed.\n
Paper
References (78)
Scroll for more · 38 remaining