RWKV: Reinventing RNNs for the Transformer Era
2023/05/22 by Bo Peng, Eric Alcaide, Peng, Bo +71 · 2 voices · 222 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Algorithm #Architecture #Artificial intelligence #Artificial neural network #Computation #Computational complexity theory #Computer engineering #Computer science #Domain Adaptation and Few-Shot Learning #Engineering #Inference #Leverage (statistics) #Machine learning #Parallelizable manifold #Quadratic growth #Recurrent neural network #Scalability #Topic Modeling #Transformer #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2305.13048
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/05/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Transformers have revolutionized almost all natural language processing (NLP) tasks but suffer from memory and computational complexity that scales quadratically with sequence length. In contrast, recurrent neural networks (RNNs) exhibit linear scaling in memory and computational requirements but struggle to match the same performance as Transformers due to limitations in parallelization and scalability. We propose a novel model architecture, Receptance Weighted Key Value (RWKV), that combines the efficient parallelizable training of transformers with the efficient inference of RNNs. Our approach leverages a linear attention mechanism and allows us to formulate the model as either a Transformer or an RNN, thus parallelizing computations during training and maintains constant computational and memory complexity during inference. We scale our models as large as 14 billion parameters, by far the largest dense RNN ever trained, and find RWKV performs on par with similarly sized Transformers, suggesting future work can leverage this architecture to create more efficient models. This work presents a significant step towards reconciling trade-offs between computational efficiency and model performance in sequence processing tasks.
Cited by
- CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
- The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory
- SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training
- Norm or Direction? Decoding Vision Mambas for High-Resolution Vision
- Convolution for Large Language Models
- The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale
- Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones
- The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
- RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways
- Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
- The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path
- Dynamics and Representation Structure of Local Approximations to Gradient-Based Learning in Linear Recurrent Neural Networks
- ASI-Evolve: AI Accelerates AI
- Effective Distillation to Hybrid xLSTM Architectures
- M2RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
- Memory Caching: RNNs with Growing Memory
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- DREMnet: An Interpretable Denoising Framework for Semi-Airborne Transient Electromagnetic Signal
- Innovative tooth segmentation using hierarchical features and bidirectional sequence modeling
- Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
- ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference
- Mercury: Ultra-Fast Language Models Based on Diffusion
- ATLAS: Learning to Optimally Memorize the Context at Test Time
- Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
- Reservoir Computing inspired Matrix Multiplication-free Language Model
- Breaking the Memory Wall: Exact Analytical Differentiation via Tiled Operator-Space Evolution
- SketchMamba: A Lightweight State-Space Model for Joint Progressive Sketch Classification and Stroke Auto-Completion
- Hierarchical Grading in Large Language Models
- ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
- UltraLBM-UNet: Ultralight Bidirectional Mamba-based Model for Skin Lesion Segmentation
- RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention
- DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis
- From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
- When Does Learning Renormalize? Sufficient Conditions for Power Law Spectral Dynamics
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- LADY: Linear Attention for Autonomous Driving Efficiency without Transformers
- SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
- Should AI Become an Intergenerational Civil Right?
- Fourier-RWKV: A Multi-State Perception Network for Efficient Image Dehazing
- Operator Lanczos Approach enabling Neural Quantum States as Real-Frequency Impurity Solvers
- GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
- FRWKV:Frequency-Domain Linear Attention for Long-Term Time Series Forecasting
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
- Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- ExplicitLM: Decoupling Knowledge from Parameters via Explicit Memory Banks
- RWKV-PCSSC: Exploring RWKV Model for Point Cloud Semantic Scene Completion
- TNT: Improving Chunkwise Training for Test-Time Memorization
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image Compression
- Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV
- UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
- Agentic World Modeling for 6G: Near-Real-Time Generative State-Space Reasoning
- Apriel-H1: Towards Efficient Enterprise Reasoning Models
- Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
- FlashEVA: Accelerating LLM inference via Efficient Attention
- Higher-order Linear Attention
- Detecting Data Contamination in LLMs via In-Context Learning
- MossNet: Mixture of State-Space Experts is a Multi-Head Attention
- Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism
- TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-shot Time Series Forecasting
- Metis: Memory Foundation Model
- Simplified Sparse Attention via Gist Tokens
- Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- ITC-RWKV: Interactive Tissue-Cell Modeling with Recurrent Key-Value Aggregation for Histopathological Subtyping
- Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
- Tapered Language Models
- AURA: Action-Gated Memory for Robot Policies at Constant VRAM
- Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
- FS-RWKV: Leveraging Frequency Spatial-Aware RWKV for 3T-to-7T MRI Translation
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL Problems
- Pack and Force Your Memory: Long-form and Consistent Video Generation
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Detecting Invariant Manifolds in ReLU-Based RNNs
- VRWKV-Editor: Reducing quadratic complexity in transformer-based video editing
- Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
- C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
- Orochi: Versatile Biomedical Image Processor
- StateX: Enhancing RNN Recall via Post-training State Expansion
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
- Beyond computational equivalence: the behavioral inference principle for machine consciousness
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
- TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
- Positional Encoding via Token-Aware Phase Attention
- Elucidating the Design Space of Decay in Linear Attention
- ACE-RL: Adaptive Constraint-Enhanced Reward for Long-form Generation Reinforcement Learning
- AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
- Quantum-Enhanced Natural Language Generation: A Multi-Model Framework with Hybrid Quantum-Classical Architectures
- ENA: Efficient N-dimensional Attention
- FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
- Hypergraph Neural Network with State Space Models for Node Classification
- Q-DPTS: Quantum Differentially Private Time Series Forecasting via Variational Quantum Circuits
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learning
- Benchmarking Quantum and Classical Sequential Models for Urban Telecommunication Forecasting
- Prototype-Driven Structure Synergy Network for Remote Sensing Images Segmentation
- Scaling Linear Attention with Sparse State Expansion
- EdgeInfinite-Instruct: Bridging SFT-Based Optimization and NPU-Level Efficiency for Edge Devices
- Onboard Hyperspectral Super-Resolution with Deep Pushbroom Neural Network
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- Flora: Effortless Context Construction to Arbitrary Length and Scale
- LowKeyEMG: Electromyographic typing with a reduced keyset
- Efficient Attention Mechanisms for Large Language Models: A Survey
- Smooth Reading: Bridging the Gap of Recurrent LLM to Self-Attention LLM on Long-Context Tasks
- When Better Eyes Lead to Blindness: A Diagnostic Study of the Information Bottleneck in CNN-LSTM Image Captioning Models
- Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding
- K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
- Synergy: End-to-end Concept Model
- On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
- U-RWKV: Lightweight medical image segmentation with direction-adaptive RWKV
- Pangenome-Informed Language Models for Synthetic Genome Sequence Generation
- Learning to Reason Across Parallel Samples for LLM Reasoning
- Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
- QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models
- InsurTech innovation using natural language processing
- SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
- If open source is to win, it must go public
- Lizard: An Efficient Linearization Framework for Large Language Models
- Scaling Attention to Very Long Sequences in Linear Time with Wavelet-Enhanced Random Spectral Attention (WERSA)
- AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling
- SAS: Simulated Attention Score
- Differential Mamba
- A Systematic Analysis of Hybrid Linear Attention
- PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
- MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection
- RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
- Memory Mosaics at scale
- Understanding and Improving Length Generalization in Recurrent Models
- MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
- AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
- Energy-Based Transformers are Scalable Learners and Thinkers
- Stem: Rethinking Causal Information Flow in Sparse Attention
- EvRWKV: A Continuous Interactive RWKV Framework for Effective Event-Guided Low-Light Image Enhancement
- VMoBA: Mixture-of-Block Attention for Video Diffusion Models
- Uncovering the Functional Roles of Nonlinearity in Memory
- Residual Matrix Transformers: Scaling the Size of the Residual Stream
- Echo State Transformer: Attention Over Finite Memories
- Accurate, fast, cheap: Choose three. Replacing Multi-Head-Attention with Bidirectional Recurrent Attention for Long-Form ASR
- Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
- Vision-QRWKV: Exploring Quantum-Enhanced RWKV Models for Image Classification
- AggTruth: Contextual Hallucination Detection using Aggregated Attention Scores in LLMs
- TPTT: Transforming Pretrained Transformers into Titans
- Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
- Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions
- Adaptable Symbolic Music Infilling with MIDI-RWKV
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
- Don't Pay Attention
- Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
- FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation
- MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
- TabFlex: Scaling Tabular Learning to Millions with Linear Attention
- Temporal Chunking Enhances Recognition of Implicit Sequential Patterns
- Channel-Imposed Fusion: A Simple yet Effective Method for Medical Time Series Classification
- SALE : Low-bit Estimation for Efficient Sparse Attention in Long-context LLM Prefilling
- TiRex: Zero-Shot Forecasting Across Long and Short Horizons with Enhanced In-Context Learning
- Scaling Reasoning without Attention
- StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
- Uncertainty Unveiled: Can Exposure to More In-context Examples Mitigate Uncertainty for Large Language Models?
- Pretraining Language Models to Ponder in Continuous Space
- Multi-View Learning with Context-Guided Receptance for Image Denoising
- Sparsified State-Space Models are Efficient Highway Networks
- Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation
- Balancing Computation Load and Representation Expressivity in Parallel Hybrid Neural Networks
- Shifting AI Efficiency From Model-Centric to Data-Centric Compression
- Do Large Language Models (Really) Need Statistical Foundations?
- Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge
- How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
- Learning What to Remember: Test-Time Training via Context Distillation
- Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
- Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN
- DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling
- Tokenization Constraints in LLMs: A Study of Symbolic and Arithmetic Reasoning Limits
- Quantum-Enhanced Channel Mixing in RWKV Models for Time Series Forecasting
- Chain-of-Model Learning for Language Model
- Bi-directional Recurrence Improves Transformer in Partially Observable Markov Decision Processes
- Concept-Guided Interpretability via Neural Chunking
- Maximizing Asynchronicity in Event-based Neural Networks
- Dyadic Mamba: Long-term Dyadic Human Motion Synthesis
- LLM4CD: Leveraging Large Language Models for Open-World Knowledge Augmented Cognitive Diagnosis
- Camera-Only 3D Panoptic Scene Completion for Autonomous Driving through Differentiable Object Shapes
- Motif-Mamba: network motif improved mamba for long-range sequence modeling
- Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environments
- FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
- Overflow Prevention Enhances Long-Context Recurrent LLMs
- Relative Overfitting and Accept-Reject Framework
- RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization
- Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
- ParalESN: Enabling parallel information processing in Reservoir Computing
- CliffordNet: All You Need is Geometric Algebra
- SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
- Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
- Nested Learning: The Illusion of Deep Learning Architectures
- R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference
- Compass: Degradation-Simulated Reciprocal Learning with Lightweight Needle RWKV for Multimodal Crack Segmentation under Missing Modalities
- Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
- Kathleen Writes: Autoregressive Generation and Data Scaling Without Attention
- Maglev: Sliding Recurrent Memory
- Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
- Event2Vec: Processing Neuromorphic Events Directly by Representations in Vector Space
- Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
- Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
- It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
- ACMamba: Fast Unsupervised Anomaly Detection via An Asymmetrical Consensus State Space Model
- The Impossibility Triangle of Long-Context Modeling
- RGB-Event based Pedestrian Attribute Recognition: A Benchmark Dataset and An Asymmetric RWKV Fusion Framework
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- SWAN-GPT: An Efficient and Scalable Approach for Long-Context Language Modeling
- Open Problems and a Hypothetical Path Forward in LLM Knowledge Paradigms
- Compound and Parallel Modes of Tropical Convolutional Neural Networks
- Lattice: Learning to Efficiently Compress the Memory
- Attention Is All You Need [wikipedia]
- History of artificial neural networks [wikipedia]
- Large language model [wikipedia]
- Transformer (deep learning) [wikipedia]
Discussions
Related