vix.ing · top · new · best · stats

Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models

2024/02/26 by Tianyi Tang, Tang, Tianyi, Wenyang Luo +13 · 39 citations
Computer Science · #Computation and Language (cs.CL) #Computer science #Computer security #FOS: Computer and information sciences #Key (lock) #Linguistics #Natural Language Processing Techniques #Natural language processing #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2402.16438

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2024/02/26 · openalex created_date 2024/02/28 · openalex updated_date 2026/07/28

Abstract

Large language models (LLMs) demonstrate remarkable multilingual capabilities without being pre-trained on specially curated multilingual parallel corpora. It remains a challenging problem to explain the underlying mechanisms by which LLMs process multilingual texts. In this paper, we delve into the composition of Transformer architectures in LLMs to pinpoint language-specific regions. Specially, we propose a novel detection method, language activation probability entropy (LAPE), to identify language-specific neurons within LLMs. Based on LAPE, we conduct comprehensive experiments on several representative LLMs, such as LLaMA-2, BLOOM, and Mistral. Our findings indicate that LLMs' proficiency in processing a particular language is predominantly due to a small subset of neurons, primarily situated in the models' top and bottom layers. Furthermore, we showcase the feasibility to "steer" the output language of LLMs by selectively activating or deactivating language-specific neurons. Our research provides important evidence to the understanding and exploration of the multilingual capabilities of LLMs.

Cited by

Related