Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
2021/08/27 by Ofir Press, Noah A. Smith, Press, Ofir +3 · 3 voices · 154 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2108.12409
Abstract
Since the introduction of the transformer model by Vaswani et al. (2017), a fundamental question has yet to be answered: how does a model achieve extrapolation at inference time for sequences that are longer than it saw during training? We first show that extrapolation can be enabled by simply changing the position representation method, though we find that current methods do not allow for efficient extrapolation. We therefore introduce a simpler and more efficient position method, Attention with Linear Biases (ALiBi). ALiBi does not add positional embeddings to word embeddings; instead, it biases query-key attention scores with a penalty that is proportional to their distance. We show that this method trains a 1.3 billion parameter model on input sequences of length 1024 that extrapolates to input sequences of length 2048, achieving the same perplexity as a sinusoidal position embedding model trained on inputs of length 2048 but training 11% faster and using 11% less memory. ALiBi's inductive bias towards recency also leads it to outperform multiple strong position methods on the WikiText-103 benchmark.
Cited by
- RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
- Pretraining Recurrent Networks without Recurrence
- Anti-Periodic Positional Encoding: Möbius Boundary Conditions Make In-Context Retrieval Reliable
- HiCI: Hierarchical Construction-Integration for Long-Context Attention
- Relative Positions Generalize, Absolute Positions Memorize: An Implicit-Bias Account of Length Generalization in Attention
- PAGE-RAG: Evidence-Grounded Adaptive Graph Retrieval for Long-Document Question Answering
- Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks
- The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture
- RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways
- 3D Geometric Tooth Alignment Planning via Deep Reinforcement Learning
- Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation
- Bifocal Attention: Harmonizing Geometric and Spectral Positional Embeddings for Algorithmic Generalization
- Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
- Long-History User Transformers for Real-Time Ad Ranking
- The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path
- RePo: Language Models with Context Re-Positioning
- Recursive Language Models
- Extending the Context of Pretrained LLMs by Dropping Their Positional Embeddings
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Decoupling the "What" and "Where" With Polar Coordinate Positional Embeddings
- Cameras as Relative Positional Encoding
- Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
- Tropical Attention: Neural Algorithmic Reasoning for Combinatorial Algorithms
- AbsenceBench: Language Models Can't Tell What's Missing
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
- Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
- Exploring Depth Generalization in Large Language Models for Solving Recursive Logic Tasks
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models
- Accelerating Language Model Workflows with Prompt Choreography
- Raven: High-Recall Sequence Modeling with Sparse Memory Routing
- Out-of-Length Scene Text Recognition: A Two-Axis Diagnosis and a Training-Free Fix
- Towards Long-window Anchoring in Vision-Language Model Distillation
- TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning
- Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
- Radiology Report Generation with Layer-Wise Anatomical Attention
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
- Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
- Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
- Rhea: Role-aware Heuristic Episodic Attention for Conversational LLMs
- FlatFormer: A Flat Transformer Knowledge Tracing Model Based on Cognitive Bias Injection
- ShaRP: SHAllow-LayeR Pruning for Video Large Language Models Acceleration
- Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
- On the Temporality for Sketch Representation Learning
- Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
- Zero-Shot Instruction Following in RL via Structured LTL Representations
- Teaching by Failure: Counter-Example-Driven Curricula for Transformer Self-Improvement
- Softmax Transformers are Turing-Complete
- UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
- Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
- Unified Spherical Frontend: Learning Rotation-Equivariant Representations of Spherical Images from Any Camera
- Selective Rotary Position Embedding
- StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion Model
- SynClaimEval: A Framework for Evaluating the Utility of Synthetic Data in Long-Context Claim Verification
- Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off
- DoPE: Denoising Rotary Position Embedding
- Generalizable Insights for Graph Transformers in Theory and Practice
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- Reduced Density Matrices Through Machine Learning
- Make It Long, Keep It Fast: End-to-End 10k-Sequence Modeling at Billion Scale on Douyin
- Retrieval Quality at Context Limit
- Hilbert-Guided Sparse Local Attention
- BudgetMem: Learning Selective Memory Policies for Cost-Efficient Long-Context Processing in Language Models
- Flashlight: PyTorch Compiler Extensions to Accelerate Attention Variants
- OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models
- Quantitative Bounds for Length Generalization in Transformers
- Spatially Grounded Concept Bottleneck Models via Part-Factorized Attention
- Journey Operators for Structured Multi-Axis Composition
- Out of Context: How important is Local Context in Neural Program Repair?
- A multimodal whole-slide foundation model for pathology
- Large Language Models as Model Organisms for Human Associative Learning
- Sparser Block-Sparse Attention via Token Permutation
- Stateful KV Cache Management for LLMs: Balancing Space, Time, Accuracy, and Positional Fidelity
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Rotate Both Ways: Time-and-Order RoPE for Generative Recommendation
- What is the Best Sequence Length for BABYLM?
- Glyph: Scaling Context Windows via Visual-Text Compression
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
- Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
- Predicting Task Performance with Context-aware Scaling Laws
- Chinese ModernBERT with Whole-Word Masking
- VideoNSA: Native Sparse Attention Scales Video Understanding
- DiffStyleTS: Diffusion Model for Style Transfer in Time Series
- PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles
- Structure Over Signal: A Globalized Approach to Multi-relational GNNs for Stock Prediction
- Forget Attention: Importance-Aware Attention Is All You Need
- Surgical Repair of Collapsed Attention Heads in ALiBi Transformers
- Design Principles for Sequence Models via Coefficient Dynamics
- Efficient Autoregressive Inference for Transformer Probabilistic Models
- Production-Grade Local LLM Inference on Apple Silicon: A Comparative Study of MLX, MLC-LLM, Ollama, llama.cpp, and PyTorch MPS
- Inverse-Free Wilson Loops for Transformers: A Practical Diagnostic for Invariance and Order Sensitivity
- Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
- Allocation of Parameters in Transformers
- Decomposing Attention To Find Context-Sensitive Neurons
- Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
- Indirect Attention: Turning Context Misalignment into a Feature
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- GenVarFormer: Predicting gene expression from long-range mutations in cancer
- Predicting Training Re-evaluation Curves Enables Effective Data Curriculums for LLMs
- Scaling with Collapse: Efficient and Predictable Training of LLM Families
- An empirical study on the limitation of Transformers in program trace generation
- Muon: Training and Trade-offs with Latent Attention and MoE
- Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
- AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
- Graph of Agents: Principled Long Context Modeling by Emergent Multi-Agent Collaboration
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- IntSR: An Integrated Generative Framework for Search and Recommendation
- LayerNorm Induces Recency Bias in Transformer Decoders
- Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization
- Synergistic Enhancement of Requirement-to-Code Traceability: A Framework Combining Large Language Model based Data Augmentation and an Advanced Encoder
- Mamba Modulation: On the Length Generalization of Mamba
- ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
- Context Recycling for Long-Horizon LLM Inference
- Transformers and genome language models
- Memory in Large Language Models: Mechanisms, Evaluation and Evolution
- Financial Risk Relation Identification through Dual-view Adaptation
- Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs
- Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential Features
- Language Modeling with Learned Meta-Tokens
- Patent Language Model Pretraining with ModernBERT
- Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
- Positional Encoding via Token-Aware Phase Attention
- Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
- CoPE: A Lightweight Complex Positional Encoding
- Improving Machine Learning-Based Robot Self-Collision Checking with Input Positional Encoding
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining
- HoPE: Hyperbolic Rotary Positional Encoding for Stable Long-Range Dependency Modeling in Large Language Models
- COGITAO: A Visual Reasoning Framework To Study Compositionality & Generalization
- SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
- Temporal social network modeling of mobile connectivity data with graph neural networks
- Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
- Efficient Large-Scale Cross-Domain Sequential Recommendation with Dynamic State Representations
- Joint Enhancement of Relational Reasoning for Long-Context LLMs
- Autoregressive Universal Video Segmentation Model
- Position Bias Mitigates Position Bias:Mitigate Position Bias Through Inter-Position Knowledge Distillation
- PENGUIN: Enhancing Transformer with Periodic-Nested Group Attention for Long-term Time Series Forecasting
- Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures
- DeepFleet: Multi-Agent Foundation Models for Mobile Robots
- Many-Turn Jailbreaking
- Human-like fleeting memory improves language learning but impairs reading time prediction in transformer language models
- Live Music Models
- Small transformer architectures for task switching
- Hidden Dynamics of Massive Activations in Transformer Training
- LaMPE: Length-aware Multi-grained Positional Encoding for Adaptive Long-context Scaling Without Training
- What are you sinking? A geometric approach on attention sink
- DELTAv2: Accelerating Dense 3D Tracking
- Context-aware Rotary Position Embedding
Discussions
Related