Current large language models (LLMs) often exhibit imbalances in multilingual capabilities and cultural adaptability, largely attributed to their English-centric pretraining data. In this paper, we introduce and investigate cross-lingual latent transplantation (<inline-formula><tex-math notation="LaTeX">$\mathcal {X}$</tex-math></inline-formula>Transplant), a probing framework which aims to further exploit the model’s internalized multilingual knowledge during inference and examine its effects on the multilingual capability and cultural adaptability of LLMs. <inline-formula><tex-math notation="LaTeX">$\mathcal {X}$</tex-math></inline-formula>Transplant framework enables models to harness the complementary strengths of both English and non-English resources by transplanting latent activations across languages. Through extensive analysis, we empirically demonstrate that <inline-formula><tex-math notation="LaTeX">$\mathcal {X}$</tex-math></inline-formula>Transplant, a form of cross-lingual interaction, has mutually beneficial effects on the multilingual capability and cultural adaptability of LLMs, particularly for low-resource languages and cultures. We further reveal that attention modules play a pivotal role in supporting multilingual understanding, while feed-forward modules are more adept at capturing culture-specific knowledge. In addition, we conduct in-depth analysis of <inline-formula><tex-math notation="LaTeX">$\mathcal {X}$</tex-math></inline-formula>Transplant’s stability, effectiveness, and generalizability. By probing the upper bound performance of <inline-formula><tex-math notation="LaTeX">$\mathcal {X}$</tex-math></inline-formula>Transplant, we expose the considerable underutilization of current LLMs’ multilingual potential—a challenge that remains open. We hope our analysis offers a new lens for advancing cross-lingual interactions and better leveraging models’ internalized multilingual knowledge.
Paper
References (65)
Scroll for more · 38 remaining