Root Mean Square Layer Normalization
2019/10/16 by Biao Zhang, Rico Sennrich, Zhang, Biao +1 · 247 citations
Computer Science · #Algorithms and Data Compression #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Neural Networks and Applications
paper · pdf · doi:10.48550/arxiv.1910.07467
openalex publication_date 2019/10/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Layer normalization (LayerNorm) has been successfully applied to various deep neural networks to help stabilize training and boost model convergence because of its capability in handling re-centering and re-scaling of both inputs and weight matrix. However, the computational overhead introduced by LayerNorm makes these improvements expensive and significantly slows the underlying network, e.g. RNN in particular. In this paper, we hypothesize that re-centering invariance in LayerNorm is dispensable and propose root mean square layer normalization, or RMSNorm. RMSNorm regularizes the summed inputs to a neuron in one layer according to root mean square (RMS), giving the model re-scaling invariance property and implicit learning rate adaptation ability. RMSNorm is computationally simpler and thus more efficient than LayerNorm. We also present partial RMSNorm, or pRMSNorm where the RMS is estimated from p% of the summed inputs without breaking the above properties. Extensive experiments on several tasks using diverse network architectures show that RMSNorm achieves comparable performance against LayerNorm but reduces the running time by 7%~64% on different models. Source code is available at https://github.com/bzhangGo/rmsnorm.
Citations
Cited by
- Towards Understanding Steering Strength
- SPECTRE: Spectral Pre-training Embeddings with Cylindrical Temporal Rotary Position Encoding for Fine-Grained sEMG-Based Movement Decoding
- Pose-Guided Residual Refinement for Interpretable Text-to-Motion Generation and Editing
- Bright 4B: Scaling Hyperspherical Learning for Segmentation in 3D Brightfield Microscopy
- The Affine Divergence: Aligning Activation Updates Beyond Normalisation
- SemCovert: Secure and Covert Video Transmission via Deep Semantic-Level Hiding
- Attention Residuals
- ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour
- PRISM: Polynomial Representations for Interaction-Structured Motor Control
- Scale Weight Decay and Train Better
- Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
- Hierarchical Grading in Large Language Models
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
- WeCon: An Efficient Weight-Conditioned Neural Solver for Multi-Objective Combinatorial Optimization Problems
- StegaFFD: Privacy-preserving Face Forgery Detection via Fine-grained Steganographic Domain Lifting
- TimeBill: Time-Budgeted Inference for Large Language Models
- Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM
- ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
- DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis
- Mitigating Forgetting in Low Rank Adaptation
- KV Admission: Learning What to Write for Efficient Long-Context Inference
- NRGPT: An Energy-based Alternative for GPT
- Sigma-MoE-Tiny Technical Report
- How Smoothing is N-simplicial Attention?
- T5Gemma 2: Seeing, Reading, and Understanding Longer
- Native and Compact Structured Latents for 3D Generation
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- Understanding and Improving Hyperbolic Deep Reinforcement Learning
- What Affects the Effective Depth of Large Language Models?
- Directional Textual Inversion for Personalized Text-to-Image Generation
- MiniLingua: A Small Open-Source LLM for European Languages
- RecTok: Reconstruction Distillation along Rectified Flow
- What Happens Next? Next Scene Prediction with a Unified Video Model
- FuXi-γ: Efficient Sequential Recommendation with Exponential-Power Temporal Encoder and Diagonal-Sparse Positional Mechanism
- CurvaDion: Curvature-Adaptive Distributed Orthonormalization
- BaRISTA: Brain Scale Informed Spatiotemporal Representation of Human Intracranial Neural Activity
- All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
- FAIR: Focused Attention Is All You Need for Generative Recommendation
- Bidirectional Normalizing Flow: From Data to Noise and Back
- Stronger Normalization-Free Transformers
- Linear socio-demographic representations emerge in Large Language Models from indirect cues
- Scaling Behavior of Discrete Diffusion Language Models
- Mixture of Lookup Key-Value Experts
- Graph Deep Learning for Intracranial Aneurysm Blood Flow Simulation and Risk Assessment
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- A scalable and real-time neural decoder for topological quantum codes
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- PCMind-2.1-Kaiyuan-2B Technical Report
- Materium: An Autoregressive Approach for Material Generation
- Multi-view Pyramid Transformer: Look Coarser to See Broader
- IE2Video: Adapting Pretrained Diffusion Models for Event-Based Video Reconstruction
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- AutoBrep: Autoregressive B-Rep Generation with Unified Topology and Geometry
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Cosine-Similarity Methods for Efficient Training and Sampling in High-Dimensional Latent Spaces
- Wukong's 72 Transformations: High-fidelity Textured 3D Morphing via Flow Models
- Subjective Depth and Timescale Transformers: Learning Where and When to Compute
- On the Origin of Algorithmic Progress in AI
- Adam Simplified: Bias Correction Debunked
- Zero-Knowledge Proof Based Verifiable Inference of Models
- MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
- PrefixGPT: Prefix Adder Optimization by a Generative Pre-trained Transformer
- Equivalence of Context and Parameter Updates in Modern Transformer Blocks
- UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification
- MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
- Predicting one-year clinical instability and mortality in heart failure patients using sequence modeling
- Decoupling Complexity from Scale in Latent Diffusion Model
- Walrus: A Cross-Domain Foundation Model for Continuum Dynamics
- ChangeDINO: DINOv3-Driven Building Change Detection in Optical Remote Sensing Imagery
- Change-of-Basis Pruning via Rotational Invariance
- Weight-sparse transformers have interpretable circuits
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
- Weaver: Kronecker Product Approximations of Spatiotemporal Attention for Traffic Network Forecasting
- CellARC: Measuring Intelligence with Cellular Automata
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Order-Level Attention Similarity Across Language Models: A Latent Commonality
- Deep Progressive Training: scaling up depth capacity of zero/one-layer models
- Motif 2 12.7B technical report
- MoSa: Motion Generation with Scalable Autoregressive Modeling
- Sample-Efficient Language Modeling with Linear Attention and Lightweight Enhancements
- The Hidden Power of Normalization Layers in Neural Networks: Exponential Capacity Control
- TetraJet-v2: Accurate NVFP4 Training for Large Language Models with Oscillation Suppression and Outlier Control
- A Sensing Whole Brain Zebrafish Foundation Model for Neuron Dynamics and Behavior
- Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
- EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities
- Running VLAs at Real-time Speed
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- Emu3.5: Native Multimodal Models are World Learners
- Angular Steering: Behavior Control via Rotation in Activation Space
- OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender
- A Unified Deep Reinforcement Learning Approach for Close Enough Traveling Salesman Problem
- Autoregressive Boltzmann Generators
- Symbol-Equivariant Recurrent Reasoning Models
- HRM-Text: Efficient Pretraining Beyond Scaling
- Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
- Generative Modeling via Drifting
- Linguistically Informed Evaluation of Multilingual ASR for African Languages
- Ministral 3
- BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
- IBNorm: Information-Bottleneck Inspired Normalization for Representation Learning
- Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling
- Language-Conditioned Representations and Mixture-of-Experts Policy for Robust Multi-Task Robotic Manipulation
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
- Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Decoder-Only Transformers
- Sequence Modeling with Spectral Mean Flows
- Beyond Higher Rank: Token-wise Input-Output Projections for Efficient Low-Rank Adaptation
- HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
- Simple Denoising Diffusion Language Models
- Toward Agents That Reason About Their Computation
- SeeDNorm: Self-Rescaled Dynamic Normalization
- Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- LongCat-Video Technical Report
- Mint: A Simple Test-Time Adaptation of Vision-Language Models against Common Corruptions
- Streaming Generation for Music Accompaniment
- Normalization in Attention Dynamics
- REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects
- A Unified Model for Multi-Task Drone Routing in Post-Disaster Road Assessment
- Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- Generative Point Tracking with Flow Matching
- The Impact of Negated Text on Hallucination with Large Language Models
- NeoDictaBERT: Pushing the Frontier of BERT models for Hebrew
- SEMPO: Lightweight Foundation Models for Time Series Forecasting
- CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
- Unifying and Enhancing Graph Transformers via a Hierarchical Mask Framework
- MoTVLA: A Vision-Language-Action Model with Unified Fast-Slow Reasoning
- JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
- Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
- VaultGemma: A Differentially Private Gemma Model
- When Embedding Models Meet: Procrustes Bounds and Applications
- Chinese ModernBERT with Whole-Word Masking
- Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
- How Well Can Preference Optimization Generalize Under Noisy Feedback?
- UniFusion: Vision-Language Model as Unified Encoder in Image Generation
- What If : Understanding Motion Through Sparse Interactions
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
- QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
- SoundReactor: Frame-level Online Video-to-Audio Generation
- Stability of Transformers under Layer Normalization
- RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation
- KORMo: Korean Open Reasoning Model for Everyone
- Scaling Laws for Code: A More Data-Hungry Regime
- Post-Norm can Resharpen Attention
- JAI-1: A Thai-Centric Large Language Model
- AWM: Accurate Weight-Matrix Fingerprint for Large Language Models
- Grouped Differential Attention
- SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba
- Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling
- Exploring the Hierarchical Reasoning Model for Small Natural-Image Classification Without Augmentation
- Flock: A Knowledge Graph Foundation Model via Learning on Random Walks
- Eliciting Chain-of-Thought Reasoning for Time Series Analysis using Reinforcement Learning
- Composer: A Search Framework for Hybrid Neural Architecture Design
- Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
- Pretraining Large Language Models with NVFP4
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
- AuON: A Linear-time Alternative to Orthogonal Momentum Updates
- UniVid: The Open-Source Unified Video Model
- Short window attention enables long-term memorization
- Training Agents Inside of Scalable World Models
- MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
- Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
- Beyond Outliers: A Study of Optimizers Under Quantization
- Memory-Efficient Fine-Tuning via Low-Rank Activation Compression
- IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
- Stochastic activations
- Multilingual Vision-Language Models, A Survey
- Compute-Optimal Quantization-Aware Training
- Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
- Real-Time Object Detection Meets DINOv3
- Probability Distribution Collapse: A Critical Bottleneck to Compact Unsupervised Neural Grammar Induction
- CCFormer: Efficient Cross-Field Interaction and Hierarchical Sequence Compression for Industrial Recommendation at Tencent
- MedLLM: An Open Medical Language Model at the Sub-Billion Scale
- mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
- ELF: Embedded Language Flows
- Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
- Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
- GWM: Towards Scalable Gaussian World Models for Robotic Manipulation
- Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
- ENSAM: an efficient foundation model for interactive segmentation of 3D medical images
- Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
- OpenViGA: Video Generation for Automotive Driving Scenes by Streamlining and Fine-Tuning Open Source Models with Public Data
- CompAir: Synergizing Complementary PIMs and In-Transit NoC Computation for Efficient LLM Acceleration
- AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
- Curriculum Learning for Mesh-based simulations
- The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
- Large Language Models Imitate Logical Reasoning, but at what Cost?
- Tenma: Robust Cross-Embodiment Robot Manipulation with Diffusion Transformer
- When FinTech Meets Privacy: Securing Financial LLMs with Differential Private Fine-Tuning
- Hyperspectral Mamba for Hyperspectral Object Tracking
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- Causal Attention with Lookahead Keys
- Dual Knowledge-Enhanced Two-Stage Reasoner for Multimodal Dialog Systems
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
- WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting
- FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
- PLaMo 2 Technical Report
- COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
- Recurrent State Encoders for Efficient Neural Combinatorial Optimization
- TRELLIS-Enhanced Surface Features for Comprehensive Intracranial Aneurysm Analysis
- ESTM: An Enhanced Dual-Branch Spectral-Temporal Mamba for Anomalous Sound Detection
- Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
- LongCat-Flash Technical Report
- QZhou-Embedding Technical Report
- MPFormer: Adaptive Framework for Industrial Multi-Task Personalized Sequential Retriever
- Provable Benefits of In-Tool Learning for Large Language Models
- IDF: Iterative Dynamic Filtering Networks for Generalizable Image Denoising
- FinCast: A Foundation Model for Financial Time-Series Forecasting
- On Surjectivity of Neural Networks: Can you elicit any behavior from your model?
- CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- Exploring Scaling Laws of CTR Model for Online Performance Improvement
- Training Transformers for Mesh-Based Simulations
- Efficient Identification of Critical Transitions via Flow Matching: A Scalable Generative Approach for Many-Body Systems
- Mamba2 Meets Silence: Robust Vocal Source Separation for Sparse Regions
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
- Maximum Score Routing For Mixture-of-Experts
- Representing Speech Through Autoregressive Prediction of Cochlear Tokens
- FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
- XQuant: Breaking the Memory Wall for LLM Inference with KV Cache Rematerialization
- A learning-driven automatic planning framework for proton PBS treatments of H&N cancers
- GBC: Generalized Behavior-Cloning Framework for Whole-Body Humanoid Imitation
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- gpt-oss-120b & gpt-oss-20b Model Card
- RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
- Multimodal learning with next-token prediction for large multimodal models
- MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows
- Matrix-Driven Identification and Reconstruction of LLM Weight Homology
- WiFo-CF: Wireless Foundation Model for CSI Feedback
- Hidden Dynamics of Massive Activations in Transformer Training
- LOST: Low-rank and Sparse Pre-training for Large Language Models
- Qwen-Image Technical Report
- Learning Dynamics of Meta-Learning in Small Model Pretraining
- MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model
- Training Dynamics of the Cooldown Stage in Warmup-Stable-Decay Learning Rate Scheduler
- Kronos: A Foundation Model for the Language of Financial Markets
- EdgeInfinite-Instruct: Bridging SFT-Based Optimization and NPU-Level Efficiency for Edge Devices
- Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
- iLRM: An Iterative Large 3D Reconstruction Model
- Improving Neural Network Training using Dynamic Learning Rate Schedule for PINNs and Image Classification
Related