As large language models get bigger, more capable, and hence more integrated into everyday human life due to wide range of applications, Dependency on the model's decisions increases, especially in more critical scenarios (e.g., medical diagnosis and financial decisions etc.), it becomes necessary to comprehend how the model functions internally in order to build trust, verify information and identify potential biases in the model's behavior. In this paper, we classify interpretability approaches based on their goals and applications while touching upon practical aspects and behavioral phenomenon discovered in recent studies.
Paper
Full text
On the Interpretability of Transformer Based Large Language Models
Semantic Scholar · 2025
Abstract
As large language models get bigger, more capable, and hence more integrated into everyday human life due to wide range of applications, Dependency on the model's decisions increases, especially in more critical scenarios (e.g., medical diagnosis and financial decisions etc.), it becomes necessary to comprehend how the model functions internally in order to build trust, verify information and identify potential biases in the model's behavior. In this paper, we classify interpretability approaches based on their goals and applications while touching upon practical aspects and behavioral phenomenon discovered in recent studies.