Beyond Language-Specific Neurons: The Challenge of Identifying Speech-Specific Neurons in Multimodal LLMs
As recent advances in multilingual large language models (LLMs) demonstrate powerful performance across numerous tasks, various studies attempt to analyze their intrinsic behavior across different languages to improve these models. Such works have expanded to the modality level, being used to detect modality-specific components (often called neurons) in the vision domain. However, it remains unclear whether such methods are also applicable to speech, another key modality used for everyday communication. In this work, we investigate whether current neuron detection methods can reliably identify neurons associated with speech processing in speech-capable LLMs. Specifically, we utilize two representative neuron detection techniques to identify candidate modality-specific neurons for speech and text, and evaluate their specialization through neuron deactivation experiments across diverse benchmarks and experimental setups. Our results show that, unlike in the text and visual modality, existing methods do not reliably detect speech-specific neurons, highlighting the limitations of current diagnostic approaches and the need for more effective methods to better interpret and improve speech LLMs.
Paper
Full text
Beyond Language-Specific Neurons: The Challenge of Identifying Speech-Specific Neurons in Multimodal LLMs
Semantic Scholar · Computer Science · 2026
Abstract
As recent advances in multilingual large language models (LLMs) demonstrate powerful performance across numerous tasks, various studies attempt to analyze their intrinsic behavior across different languages to improve these models. Such works have expanded to the modality level, being used to detect modality-specific components (often called neurons) in the vision domain. However, it remains unclear whether such methods are also applicable to speech, another key modality used for everyday communication. In this work, we investigate whether current neuron detection methods can reliably identify neurons associated with speech processing in speech-capable LLMs. Specifically, we utilize two representative neuron detection techniques to identify candidate modality-specific neurons for speech and text, and evaluate their specialization through neuron deactivation experiments across diverse benchmarks and experimental setups. Our results show that, unlike in the text and visual modality, existing methods do not reliably detect speech-specific neurons, highlighting the limitations of current diagnostic approaches and the need for more effective methods to better interpret and improve speech LLMs.