Interpeval: A Multi-Faceted Framework for Assessing Interpretability in Large Language Models

Large language models (LLMs) have demonstrated extraordinary capabilities across a wide range of tasks, achieving near-human or superhuman performance. However, their powerful abilities also raise concerns about trustworthiness and misuse, making interpretability a critical factor for ensuring their reliability and control. While existing research has provided tools and insights into understanding LLMs' mechanisms, the question of whether one model is inherently more interpretable than another remains largely unexplored. To address this gap, we propose Interpeval, a comprehensive evaluation framework for model interpretability across three dimensions: parametric, functional, and behavioral interpretability, encompassing nine detailed metrics. Using InterpEval, we assessed ten diverse and widely used open-source LLMs. Our findings reveal that scaling the model size and applying instruction tuning was shown to improve interpretability, offering valuable guidance for the development of more interpretable models. This work provides a new benchmark for evaluating interpretability and actionable insights for advancing trustworthy AI.

Paper

Full text

PDF

Interpeval: A Multi-Faceted Framework for Assessing Interpretability in Large Language Models

Semantic Scholar · 2025

Abstract

Large language models (LLMs) have demonstrated extraordinary capabilities across a wide range of tasks, achieving near-human or superhuman performance. However, their powerful abilities also raise concerns about trustworthiness and misuse, making interpretability a critical factor for ensuring their reliability and control. While existing research has provided tools and insights into understanding LLMs' mechanisms, the question of whether one model is inherently more interpretable than another remains largely unexplored. To address this gap, we propose Interpeval, a comprehensive evaluation framework for model interpretability across three dimensions: parametric, functional, and behavioral interpretability, encompassing nine detailed metrics. Using InterpEval, we assessed ten diverse and widely used open-source LLMs. Our findings reveal that scaling the model size and applying instruction tuning was shown to improve interpretability, offering valuable guidance for the development of more interpretable models. This work provides a new benchmark for evaluating interpretability and actionable insights for advancing trustworthy AI.

Similar papers

© 2026 NYSGPT2525 LLC