LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
2024/02/21 by Yiran Ding, Ding, Yiran, Li Lyna Zhang +13 · 3 voices · 98 citations
Computer Science · #Archaeology #Computer science #Context (archaeology) #History #Image Processing and 3D Reconstruction #Real-time computing #Window (computing) #Window of opportunity #World Wide Web #cs.CL
paper · pdf · doi:10.48550/arxiv.2402.13753
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/02/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large context window is a desirable feature in large language models (LLMs). However, due to high fine-tuning costs, scarcity of long texts, and catastrophic values introduced by new token positions, current extended context windows are limited to around 128k tokens. This paper introduces LongRoPE that, for the first time, extends the context window of pre-trained LLMs to an impressive 2048k tokens, with up to only 1k fine-tuning steps at within 256k training lengths, while maintaining performance at the original short context window. This is achieved by three key innovations: (i) we identify and exploit two forms of non-uniformities in positional interpolation through an efficient search, providing a better initialization for fine-tuning and enabling an 8x extension in non-fine-tuning scenarios; (ii) we introduce a progressive extension strategy that first fine-tunes a 256k length LLM and then conducts a second positional interpolation on the fine-tuned extended LLM to achieve a 2048k context window; (iii) we readjust LongRoPE on 8k length to recover the short context window performance. Extensive experiments on LLaMA2 and Mistral across various tasks demonstrate the effectiveness of our method. Models extended via LongRoPE retain the original architecture with minor modifications to the positional embedding, and can reuse most pre-existing optimizations.
Cited by
- Anti-Periodic Positional Encoding: Möbius Boundary Conditions Make In-Context Retrieval Reliable
- Harmonia: Algorithm-Hardware Co-Design for Memory- and Compute-Efficient BFP-based LLM Inference
- HiCI: Hierarchical Construction-Integration for Long-Context Attention
- RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
- StreamingVLM: Real-Time Understanding for Infinite Video Streams
- Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings
- EntropyLong: Effective Long-Context Training via Predictive Uncertainty
- AstraNav-Memory: Contexts Compression for Long Memory
- Towards Long-window Anchoring in Vision-Language Model Distillation
- Decoupling Positional and Symbolic Attention Behavior in Transformers
- Spatia: Video Generation with Updatable Spatial Memory
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
- Adaptive Soft Rolling KV Freeze with Entropy-Guided Recovery: Sublinear Memory Growth for Efficient LLM Inference
- Persian-Phi: Efficient Cross-Lingual Adaptation of Compact LLMs via Curriculum Learning
- Rhea: Role-aware Heuristic Episodic Attention for Conversational LLMs
- ChipMind: Retrieval-Augmented Reasoning for Long-Context Circuit Design Specifications
- Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
- InnoGym: Benchmarking the Innovation Potential of AI Agents
- Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis
- Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
- DoPE: Denoising Rotary Position Embedding
- Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
- Iterative Critique-Refine Framework for Enhancing LLM Personalization
- Tagging-Augmented Generation: Assisting Language Models in Finding Intricate Knowledge In Long Contexts
- LooGLE v2: Are LLMs Ready for Real World Long Dependency Challenges?
- LoongRL: Reinforcement Learning for Advanced Reasoning over Long Contexts
- MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
- Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
- NOSA: Native and Offloadable Sparse Attention
- A Survey on Agentic Multimodal Large Language Models
- COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
- Inverse-Free Wilson Loops for Transformers: A Practical Diagnostic for Invariance and Order Sensitivity
- Mid-Training of Large Language Models: A Survey
- Revisiting Long-context Modeling from Context Denoising Perspective
- Multilingual Vision-Language Models, A Survey
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
- SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers
- WolBanking77: Wolof Banking Speech Intent Classification Dataset
- Positional Encoding via Token-Aware Phase Attention
- Length-Aware Rotary Position Embedding for Text-Speech Alignment
- Evalet: Evaluating Large Language Models through Functional Fragmentation
- CCF: A Context Compression Framework for Efficient Long-Sequence Language Modeling
- Tree of Agents: Improving Long-Context Capabilities of Large Language Models through Multi-Perspective Reasoning
- ACE-RL: Adaptive Constraint-Enhanced Reward for Long-form Generation Reinforcement Learning
- Memory Limitations of Prompt Tuning in Transformers
- Addressing accuracy and hallucination of LLMs in Alzheimer's disease research through knowledge graphs
- Joint Enhancement of Relational Reasoning for Long-Context LLMs
- Towards a Holistic and Automated Evaluation Framework for Multi-Level Comprehension of LLMs in Book-Length Contexts
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- The Roots of International Perceptions: Simulating US Attitude Changes Towards China with LLM Agents
- Exploring the Challenges and Opportunities of AI-assisted Codebase Generation
- Key-Augmented Neural Triggers for Knowledge Sharing
- LaMPE: Length-aware Multi-grained Positional Encoding for Adaptive Long-context Scaling Without Training
- NeedleChain: Measuring Intact Long-Context Reasoning Capability of Large Language Models
- Flora: Effortless Context Construction to Arbitrary Length and Scale
- Routine: A Structural Planning Framework for LLM Agent System in Enterprise
- Towards Compute-Optimal Many-Shot In-Context Learning
- Extrapolation by Association: Length Generalization Transfer in Transformers
- Pimba: A Processing-in-Memory Acceleration for Post-Transformer Large Language Model Serving
- To Trade or Not to Trade: An Agentic Approach to Estimating Market Risk Improves Trading Decisions
- SAS: Simulated Attention Score
- Understanding and Improving Length Generalization in Recurrent Models
- Edit Flows: Flow Matching with Edit Operations
- Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
- CommVQ: Commutative Vector Quantization for KV Cache Compression
- MiniCPM4: Ultra-Efficient LLMs on End Devices
- PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding
- LongLLaDA: Unlocking Long Context Capabilities in Diffusion LLMs
- Lag-Relative Sparse Attention In Long Context Training
- The NTNU System at the S&I Challenge 2025 SLA Open Track
- Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching
- TextVidBench: A Benchmark for Long Video Scene Text Understanding
- Rethinking LLM Advancement: Compute-Dependent and Independent Paths to Progress
- Native-Resolution Image Synthesis
- Evaluating the Sensitivity of LLMs to Prior Context
- What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
- Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
- Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
- Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning
- Same Semantics, Different Paths: Self-Improving Alignment for Vision-Text Compression
- NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
- Deconstructing Positional Information: From Attention Logits to Training Biases
- PSC: Extending Context Window of Large Language Models via Phase Shift Calibration
- Parallel Scaling Law for Language Models
- LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis
- Scaling Context, Not Parameters: Training a Compact 7B Language Model for Efficient Long-Context Processing
- FreqKV: Frequency Domain Key-Value Compression for Efficient Context Window Extension
- Safety in Batches? Understanding and Mitigating Safety Failures in Batch Prompting
- ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
- KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
- Screening Is Enough
- Reconstructing Context: Evaluating Advanced Chunking Strategies for Retrieval-Augmented Generation
- Generative AI Literacy: A Comprehensive Framework for Literacy and Responsible Use
- Effective Length Extrapolation via Dimension-Wise Positional Embeddings Manipulation
- Sensitivity-Positional Co-Localization in GQA Transformers
Discussions
Related