Retentive Network: A Successor to Transformer for Large Language Models
2023/07/17 by Yutao Sun, Li Dong, Sun, Yutao +13 · 8 voices · 191 citations
Computer Science · #Algorithm #Artificial intelligence #Computer science #Decoding methods #Inference #Language model #Metamodeling #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Programming language #Topic Modeling #Transformer
paper · pdf · doi:10.48550/arxiv.2307.08621
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/07/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
In this work, we propose Retentive Network (RetNet) as a foundation architecture for large language models, simultaneously achieving training parallelism, low-cost inference, and good performance. We theoretically derive the connection between recurrence and attention. Then we propose the retention mechanism for sequence modeling, which supports three computation paradigms, i.e., parallel, recurrent, and chunkwise recurrent. Specifically, the parallel representation allows for training parallelism. The recurrent representation enables low-cost O(1) inference, which improves decoding throughput, latency, and GPU memory without sacrificing performance. The chunkwise recurrent representation facilitates efficient long-sequence modeling with linear complexity, where each chunk is encoded parallelly while recurrently summarizing the chunks. Experimental results on language modeling show that RetNet achieves favorable scaling results, parallel training, low-cost deployment, and efficient inference. The intriguing properties make RetNet a strong successor to Transformer for large language models. Code will be available at https://aka.ms/retnet.
Cited by
- Pretraining Recurrent Networks without Recurrence
- The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale
- Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones
- T2MLR: Transformer with Temporal Middle-Layer Recurrence
- Concept-Guided Spatial Regularization for World Models in Atari Pong
- Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation
- Online Neural Space Time Memory for Dynamic Novel View Synthesis
- Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions
- LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning
- JAXBench: Benchmarking Autonomous TPU Kernel Optimization
- Effective Distillation to Hybrid xLSTM Architectures
- Mamba-3: Improved Sequence Modeling using State Space Principles
- M2RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
- Memory Caching: RNNs with Growing Memory
- Kimi Linear: An Expressive, Efficient Attention Architecture
- ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference
- Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
- Fast weight programming and linear transformers: from machine learning to neurobiology
- Blending Complementary Memory Systems in Hybrid Quadratic-Linear Transformers
- ATLAS: Learning to Optimally Memorize the Context at Test Time
- Log-Linear Attention
- SPECTRE: An FFT-Based Efficient Drop-In Replacement to Self-Attention for Long Contexts
- Breaking the Memory Wall: Exact Analytical Differentiation via Tiled Operator-Space Evolution
- Attention Residuals
- SketchMamba: A Lightweight State-Space Model for Joint Progressive Sketch Classification and Stroke Auto-Completion
- Memory for Large Language Models
- Raven: High-Recall Sequence Modeling with Sparse Memory Routing
- Renormalization-Group Geometry of Homeostatically Regulated Reentry Networks
- DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- LADY: Linear Attention for Autonomous Driving Efficiency without Transformers
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- GDKVM: Echocardiography Video Segmentation via Spatiotemporal Key-Value Memory with Gated Delta Rule
- SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processing
- Diffusion Is Your Friend in Show, Suggest and Tell
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
- Continuous-Time Homeostatic Dynamics for Reentrant Inference Models
- Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
- Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
- Selective Rotary Position Embedding
- Goal-Directed Search Outperforms Goal-Agnostic Memory Compression in Long-Context Memory Tasks
- CAMS: Towards Compositional Zero-Shot Learning via Gated Cross-Attention and Multi-Space Disentanglement
- Dynamic Nested Hierarchies: Pioneering Self-Evolution in Machine Learning Architectures for Lifelong Intelligence
- TNT: Improving Chunkwise Training for Test-Time Memorization
- Recursive Dynamics in Fast-Weights Homeostatic Reentry Networks: Toward Reflective Intelligence
- UniLION: Towards Unified Autonomous Driving Model with Linear Group RNNs
- Agentic World Modeling for 6G: Near-Real-Time Generative State-Space Reasoning
- Apriel-H1: Towards Efficient Enterprise Reasoning Models
- Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
- FlashEVA: Accelerating LLM inference via Efficient Attention
- Higher-order Linear Attention
- Nirvana: A Specialized Generalist Model With Task-Aware Memory Mechanism
- Metis: Memory Foundation Model
- Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
- Simplified Sparse Attention via Gist Tokens
- Long-Context Modeling with Dynamic Hierarchical Sparse Attention for On-Device LLMs
- Alias-Free ViT: Fractional Shift Invariance via Linear Attention
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- Stateful KV Cache Management for LLMs: Balancing Space, Time, Accuracy, and Positional Fidelity
- Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
- Chimera: State Space Models Beyond Sequences
- Tapered Language Models
- AURA: Action-Gated Memory for Robot Policies at Constant VRAM
- Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction
- Attention Once Is All You Need: Efficient Streaming Inference with Stateful Transformers
- HeSRN: Representation Learning On Heterogeneous Graphs via Slot-Aware Retentive Network
- Design Principles for Sequence Models via Coefficient Dynamics
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- Recurrence-Complete Frame-based Action Models
- Untangling Component Imbalance in Hybrid Linear Attention Conversion Methods
- On Structured State-Space Duality
- Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
- An Early Exploration of Deep-Learning-Driven Prefetching for Far Memory
- Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
- Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
- VRWKV-Editor: Reducing quadratic complexity in transformer-based video editing
- TTT3R: 3D Reconstruction as Test-Time Training
- Context-Driven Performance Modeling for Causal Inference Operators on Neural Processing Units
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention
- Enhancing Linear Attention with Residual Learning
- SeedPolicy: Horizon Scaling via Self-Evolving Diffusion Policy for Robot Manipulation
- Transformers and genome language models
- Beyond computational equivalence: the behavioral inference principle for machine consciousness
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Achilles' Heel of Mamba: Essential difficulties of the Mamba architecture demonstrated by synthetic data
- SAGA: Selective Adaptive Gating for Efficient and Expressive Linear Attention
- Large Language Model Scaling Laws for Neural Quantum States in Quantum Chemistry
- Point-Plane Projections for Accurate LiDAR Semantic Segmentation in Small Data Scenarios
- Elucidating the Design Space of Decay in Linear Attention
- SpikingBrain: Spiking Brain-inspired Large Models
- Differentiable Entropy Regularization: A Complexity-Aware Approach for Neural Optimization
- AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
- On Surjectivity of Neural Networks: Can you elicit any behavior from your model?
- The Computational Complexity of Satisfiability in State Space Models
- ENA: Efficient N-dimensional Attention
- Scaling Linear Attention with Sparse State Expansion
- Understanding Transformers through the Lens of Pavlovian Conditioning
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- Efficient Attention Mechanisms for Large Language Models: A Survey
- MeMo: Memory as a Model
- Learning State-Tracking from Code Using Linear RNNs
- On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention
- From Matching to Generation: A Survey on Generative Information Retrieval
- SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
- Hybrid Systolic Array Accelerator with Optimized Dataflow for Edge Large Language Model Inference
- SAS: Simulated Attention Score
- A Survey on Latent Reasoning
- A Systematic Analysis of Hybrid Linear Attention
- MambaFusion: Height-Fidelity Dense Global Fusion for Multi-modal 3D Object Detection
- Scaling Context Requires Rethinking Attention
- RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
- Understanding and Improving Length Generalization in Recurrent Models
- ZeCO: Zero Communication Overhead Sequence Parallelism for Linear Attention
- Residual Matrix Transformers: Scaling the Size of the Residual Stream
- Norm×Direction: Restoring the Missing Query Norm in Vision Linear Attention
- Echo State Transformer: Attention Over Finite Memories
- Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
- Polyline Path Masked Attention for Vision Transformer
- T-SHRED: Symbolic Regression for Regularization and Model Discovery with Transformer Shallow Recurrent Decoders
- RATTENTION: Towards the Minimal Sliding Window Size in Local-Global Attention Models
- A Gravity-informed Spatiotemporal Transformer for Human Activity Intensity Prediction
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- HyRet-Change: A hybrid retentive network for remote sensing change detection
- Don't Pay Attention
- Sequential-Parallel Duality in Prefix Scannable Models
- Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
- MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
- TabFlex: Scaling Tabular Learning to Millions with Linear Attention
- Comba: Improving Bilinear RNNs with Closed-loop Control
- A New Spatiotemporal Correlation Anomaly Detection Method that Integrates Contrastive Learning and Few-Shot Learning in Wireless Sensor Networks
- Learning Semantics, Not Addresses: Runtime Neural Prefetching for Far Memory
- Test-Time Training Done Right
- LoLA: Low-Rank Linear Attention With Sparse Caching
- Scaling Reasoning without Attention
- Oryx: a Scalable Sequence Model for Many-Agent Coordination in Offline MARL
- Curse of High Dimensionality Issue in Transformer for Long-context Modeling
- StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
- HAD: Hybrid Architecture Distillation Outperforms Teacher in Genomic Sequence Modeling
- RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
- Sparsified State-Space Models are Efficient Highway Networks
- Understanding Transformer from the Perspective of Associative Memory
- Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge
- FAR: Function-preserving Attention Replacement for IMC-friendly Inference
- How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
- Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse
- DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling
- MultiSenseSeg: A Cost-Effective Unified Multimodal Semantic Segmentation Model for Remote Sensing
- Parallel Layer Normalization for Universal Approximation
- Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform
- Maximizing Asynchronicity in Event-based Neural Networks
- Motif-Mamba: network motif improved mamba for long-range sequence modeling
- Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environments
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
- The Transformer as a Polar State Estimator
- LLM Security: Vulnerabilities, Attacks, Defenses, and Countermeasures
- Time is Not Compute: Scaling Laws for Wall-Clock Constrained Training on Consumer GPUs
- FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
- In-Place Test-Time Training
- Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
- SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
- Graph Fourier Transformer with Structure-Frequency Information
- Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling
- The Key to Going Linear: Analysis-Driven Transformer Linearization
- Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
- Screening Is Enough
- Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
- Nested Learning: The Illusion of Deep Learning Architectures
- WuNeng: Hybrid State with Attention
- State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking
- Attention Needs to Focus: A Unified Perspective on Attention Allocation
- Unveiling the Hidden: Movie Genre and User Bias in Spoiler Detection
- Maglev: Sliding Recurrent Memory
- Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
- Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
- Event2Vec: Processing Neuromorphic Events Directly by Representations in Vector Space
- Hadamard product in deep learning: Introduction, Advances and Challenges
- It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
- Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers
- The Impossibility Triangle of Long-Context Modeling
- Millions of States: Designing a Scalable MoE Architecture with RWKV-7 Meta-learner
- Compound and Parallel Modes of Tropical Convolutional Neural Networks
- Multihead self-attention in cortico-thalamic circuits
- Lattice: Learning to Efficiently Compress the Memory
- DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
Discussions
- Retentive Network: A Successor to Transformer for Large Language Models [hn, 112 points, 19 comments]
- Retentive Network: A Successor to Transformer for Large Language Models [hn, 11 points, 3 comments]
- Retentive Network: A Successor to Transformer for Large Language Models [hn, 6 points, 0 comments]
- New research presents the Retentive Network (RetNet), a potential successor to Transformers, offering parallel training, cost-efficient inference, and effective large-scale language modeling. #AI #Mac [bsky, 5 points, 0 comments]
- Retentive Network: A Successor to Transformer for Large Language Models [hn, 4 points, 1 comments]
- Retentive Network: A Successor to Transformer for Large Language Models [hn, 3 points, 0 comments]
- Retentive Network: A Successor to Transformer for Large Language Models https://arxiv.org/abs/2307.08621 テクニカルには面白いけどタイトルは好きじゃない.これが導入しているのは Transformer の successor ではなく ResNet の variant.時系列に対する AR モ [bsky, 2 points, 1 comments]
- これかTransformer越えと噂のモデルの論文は / "Retentive Network: A Successor to Transformer for Large Language Models "https://arxiv.org/abs/2307.08621 [bsky, 1 points, 0 comments]
Related