Evaluating the Interpretability of Generative Models by Interactive Reconstruction

For machine learning models to be most useful in numerous sociotechnical\nsystems, many have argued that they must be human-interpretable. However,\ndespite increasing interest in interpretability, there remains no firm\nconsensus on how to measure it. This is especially true in representation\nlearning, where interpretability research has focused on "disentanglement"\nmeasures only applicable to synthetic datasets and not grounded in human\nfactors. We introduce a task to quantify the human-interpretability of\ngenerative model representations, where users interactively modify\nrepresentations to reconstruct target instances. On synthetic datasets, we find\nperformance on this task much more reliably differentiates entangled and\ndisentangled models than baseline approaches. On a real dataset, we find it\ndifferentiates between representation learning methods widely believed but\nnever shown to produce more or less interpretable models. In both cases, we ran\nsmall-scale think-aloud studies and large-scale experiments on Amazon\nMechanical Turk to confirm that our qualitative and quantitative results\nagreed.\n

Paper

References (78)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC