Ring Attention with Blockwise Transformers for Near-Infinite Context
2023/10/03 by Hao Liu, Matei Zaharia, Liu, Hao +3 · 1 voice · 95 citations
Computer Science · #Topic Modeling #Domain Adaptation and Few-Shot Learning #Advanced Neural Network Applications
paper · pdf · doi:10.48550/arxiv.2310.01889
Abstract
Transformers have emerged as the architecture of choice for many state-of-the-art AI models, showcasing exceptional performance across a wide range of AI applications. However, the memory demands imposed by Transformers limit their ability to handle long sequences, thereby posing challenges in utilizing videos, actions, and other long-form sequences and modalities in complex environments. We present a novel approach, Ring Attention with Blockwise Transformers (Ring Attention), which leverages blockwise computation of self-attention and feedforward to distribute long sequences across multiple devices while fully overlapping the communication of key-value blocks with the computation of blockwise attention. Our approach enables training and inference of sequences that are up to device count times longer than those achievable by prior memory-efficient Transformers, without resorting to approximations or incurring additional communication and computation overheads. Extensive experiments on language modeling and reinforcement learning tasks demonstrate the effectiveness of our approach in allowing millions of tokens context size and improving performance.
Cited by
- LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
- A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix
- DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
- FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
- Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
- Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
- InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context
- The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path
- Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention
- Training Foundation Models on a Full-Stack AMD Platform: Compute, Networking, and System Design
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
- AbsenceBench: Language Models Can't Tell What's Missing
- Log-Linear Attention
- Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
- StreamIndex: Memory-Bounded Compressed Sparse Attention via Streaming Top-k
- TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
- CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention
- Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
- Designing Spatial Architectures for Sparse Attention: STAR Accelerator via Cross-Stage Tiling
- Memory as Resonance: A Biomimetic Architecture for Infinite Context Memory on Ergodic Phonetic Manifolds
- RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
- Shuttling Compiler for Trapped-Ion Quantum Computers Based on Large Language Models
- Kling-Omni Technical Report
- INTELLECT-3: Technical Report
- LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
- Spatia: Video Generation with Updatable Spatial Memory
- TEMP: A Memory Efficient Physical-aware Tensor Partition-Mapping Framework on Wafer-scale Chips
- Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
- Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
- GoodSpeed: Optimizing Fair Goodput with Adaptive Speculative Decoding in Distributed Edge Inference
- Rhea: Role-aware Heuristic Episodic Attention for Conversational LLMs
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- db-SP: Accelerating Sparse Attention for Visual Generative Models with Dual-Balanced Sequence Parallelism
- HTTM: Head-wise Temporal Token Merging for Faster VGGT
- Block Cascading: Training Free Acceleration of Block-Causal Video Models
- Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Pier: Efficient Large Language Model pretraining with Relaxed Global Communication
- Optimizing Mixture of Block Attention
- Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
- MoFa: A Unified Performance Modeling Framework for LLM Pretraining
- π-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
- StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
- Optimizing Long-context LLM Serving via Fine-grained Sequence Parallelism
- Hilbert-Guided Sparse Local Attention
- FlashEVA: Accelerating LLM inference via Efficient Attention
- TridentServe: A Stage-level Serving System for Diffusion Pipelines
- LongCat-Video Technical Report
- Sparser Block-Sparse Attention via Token Permutation
- AsyncHZP: Hierarchical ZeRO Parallelism with Asynchronous Scheduling for Scalable LLM Training
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
- MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
- Efficient Long-context Language Model Training by Core Attention Disaggregation
- On Pretraining for Project-Level Code Completion
- DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
- Haystack Engineering: Context Engineering for Heterogeneous and Agentic Long-Context Evaluation
- ELMUR: External Layer Memory with Update/Rewrite for Long-Horizon RL Problems
- LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling
- Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap
- TetriServe: Efficiently Serving Mixed DiT Workloads
- Revisiting Long-context Modeling from Context Denoising Perspective
- Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space
- Toward Co-adapting Machine Learning Job Shape and Cluster Topology
- TASP: Topology-aware Sequence Parallelism
- SlimPack: Fine-Grained Asymmetric Packing for Balanced and Efficient Variable-Length LLM Training
- Parallax: Efficient LLM Inference Service over Decentralized Environment
- A Scalable Distributed Framework for Multimodal GigaVoxel Image Registration
- ResFormer: All-Time Reservoir Memory for Long Sequence Classification
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
- Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
- Data-Centric Elastic Pipeline Parallelism for Efficient Long-Context LLM Training
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
- RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
- BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
- Mamba Modulation: On the Length Generalization of Mamba
- ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate
- Training Compute-Optimal Protein Language Models
- An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
- Robust LLM Training Infrastructure at ByteDance
- SuperGen: An Efficient Ultra-high-resolution Video Generation System with Sketching and Tiling
- AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions
- TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving
- Learned Structure in Cartridges: Keys as Shareable Routers in Self-Studied Representations
- VARCO-VISION-2.0 Technical Report
- veScale: Consistent and Efficient Tensor Programming with Eager-Mode SPMD
- MeVe: A Modular System for Memory Verification and Effective Context Control in Language Models
- Chunked TabPFN: Exact Training-Free In-Context Learning for Long-Context Tabular Data
- Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
- WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
- Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context
- KnapFormer: An Online Load Balancer for Efficient Diffusion Transformers Training
- PiKV: KV Cache Management System for Mixture of Experts
- G-Core: A Simple, Scalable and Balanced RLHF Trainer
Discussions
Related