On the Interpretability of Transformer Based Large Language Models

As large language models get bigger, more capable, and hence more integrated into everyday human life due to wide range of applications, Dependency on the model's decisions increases, especially in more critical scenarios (e.g., medical diagnosis and financial decisions etc.), it becomes necessary to comprehend how the model functions internally in order to build trust, verify information and identify potential biases in the model's behavior. In this paper, we classify interpretability approaches based on their goals and applications while touching upon practical aspects and behavioral phenomenon discovered in recent studies.

Paper

Full text

PDF

On the Interpretability of Transformer Based Large Language Models

Semantic Scholar · 2025

Abstract

As large language models get bigger, more capable, and hence more integrated into everyday human life due to wide range of applications, Dependency on the model's decisions increases, especially in more critical scenarios (e.g., medical diagnosis and financial decisions etc.), it becomes necessary to comprehend how the model functions internally in order to build trust, verify information and identify potential biases in the model's behavior. In this paper, we classify interpretability approaches based on their goals and applications while touching upon practical aspects and behavioral phenomenon discovered in recent studies.

Similar papers

© 2026 NYSGPT2525 LLC