The Achilles' Heel of LLMs: How Altering a Handful of Neurons Can Cripple Language Abilities
2025/10/11 by Qin, Zixuan, Lyu, Kunlin, Yu, Qingchen +2 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.10238
Abstract
Large Language Models (LLMs) have become foundational tools in natural language processing, powering a wide range of applications and research. Many studies have shown that LLMs share significant similarities with the human brain. Recent neuroscience research has found that a small subset of biological neurons in the human brain are crucial for core cognitive functions, which raises a fundamental question: do LLMs also contain a small subset of critical neurons? In this paper, we investigate this question by proposing a Perturbation-based Causal Identification of Critical Neurons method to systematically locate such critical neurons in LLMs. Our findings reveal three key insights: (1) LLMs contain ultra-sparse critical neuron sets. Disrupting these critical neurons can cause a 72B-parameter model with over 1.1 billion neurons to completely collapse, with perplexity increasing by up to 20 orders of magnitude; (2) These critical neurons are not uniformly distributed, but tend to concentrate in the outer layers, particularly within the MLP down_proj components; (3) Performance degradation exhibits sharp phase transitions, rather than a gradual decline, when these critical neurons are disrupted. Through comprehensive experiments across diverse model architectures and scales, we provide deeper analysis of these phenomena and their implications for LLM robustness and interpretability. These findings can offer guidance for developing more robust model architectures and improving deployment security in safety-critical applications.
Citations
- Negative Pre-activations Differentiate Syntax
- GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
- SAND-Math: Using LLMs to Generate Novel, Difficult and Useful Mathematics Questions and Answers
- Vision Transformers Don't Need Trained Registers
- Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning
- A Closer Look at Multimodal Representation Collapse
- Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence
- NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models
- Why do LLMs attend to the first token?
- Interpreting the Repeated Token Phenomenon in Large Language Models
- Visual Reasoning Evaluation of Grok, Deepseek Janus, Gemini, Qwen, Mistral, and ChatGPT
- Computational functions of precisely balanced neuronal microcircuits in an olfactory memory network
- The Super Weight in Large Language Models
- Measuring short-form factuality in large language models
- Large Language Models and Cognitive Science: A Comprehensive Review of Similarities, Differences, and Challenges
- Transformer Layers as Painters
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
- Learnable Privacy Neurons Localization in Language Models
- ReconBoost: Boosting Can Achieve Modality Reconcilement
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Massive Activations in Large Language Models
- HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
- Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models
- Neuron-based Multifractal Analysis of Neuron Interaction Dynamics in Large Models
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Instruction-Following Evaluation for Large Language Models
- Efficient Streaming Language Models with Attention Sinks
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
- PMET: Precise Model Editing in a Transformer
- DenseMP: Unsupervised Dense Pre-training for Few-shot Medical Image Segmentation
- Large Language Models
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model
- Language Models are Multilingual Chain-of-Thought Reasoners
- Locating and Editing Factual Associations in GPT
- Understanding and Overcoming the Challenges of Efficient Transformer Quantization
- A Plug-and-Play Method for Controlled Text Generation
- BERT Busters: Outlier Dimensions that Disrupt Transformers
- Knowledge Neurons in Pretrained Transformers
- Measuring Mathematical Problem Solving With the MATH Dataset
- On the Dangers of Stochastic Parrots
- Towards Robust Neural Networks via Close-loop Control
- Provably-Robust Runtime Monitoring of Neuron Activation Patterns
- On sparse connectivity, adversarial robustness, and a novel model of the\n artificial neuron
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Runtime Monitoring Neuron Activation Patterns
Cited by
Related