Interpeval: A Multi-Faceted Framework for Assessing Interpretability in Large Language Models
Large language models (LLMs) have demonstrated extraordinary capabilities across a wide range of tasks, achieving near-human or superhuman performance. However, their powerful abilities also raise concerns about trustworthiness and misuse, making interpretability a critical factor for ensuring their reliability and control. While existing research has provided tools and insights into understanding LLMs' mechanisms, the question of whether one model is inherently more interpretable than another remains largely unexplored. To address this gap, we propose Interpeval, a comprehensive evaluation framework for model interpretability across three dimensions: parametric, functional, and behavioral interpretability, encompassing nine detailed metrics. Using InterpEval, we assessed ten diverse and widely used open-source LLMs. Our findings reveal that scaling the model size and applying instruction tuning was shown to improve interpretability, offering valuable guidance for the development of more interpretable models. This work provides a new benchmark for evaluating interpretability and actionable insights for advancing trustworthy AI.
Paper
Full text
Interpeval: A Multi-Faceted Framework for Assessing Interpretability in Large Language Models
Semantic Scholar · 2025
Abstract
Large language models (LLMs) have demonstrated extraordinary capabilities across a wide range of tasks, achieving near-human or superhuman performance. However, their powerful abilities also raise concerns about trustworthiness and misuse, making interpretability a critical factor for ensuring their reliability and control. While existing research has provided tools and insights into understanding LLMs' mechanisms, the question of whether one model is inherently more interpretable than another remains largely unexplored. To address this gap, we propose Interpeval, a comprehensive evaluation framework for model interpretability across three dimensions: parametric, functional, and behavioral interpretability, encompassing nine detailed metrics. Using InterpEval, we assessed ten diverse and widely used open-source LLMs. Our findings reveal that scaling the model size and applying instruction tuning was shown to improve interpretability, offering valuable guidance for the development of more interpretable models. This work provides a new benchmark for evaluating interpretability and actionable insights for advancing trustworthy AI.