vix.ing · top · new · best · stats · spec

Critical attention scaling in long-context transformers

2025/10/07 by Shi Chen, Zhengjiang Lin, Chen, Shi +5 · 2 citations
Computer Science · Mathematics · #Artificial Intelligence (cs.AI) #Classical Analysis and ODEs (math.CA) #Discrete Mathematics (cs.DM) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #cs.AI #cs.DM #cs.LG #math.CA

paper · pdf · doi:10.48550/arxiv.2510.05554

published as Proceedings of the Fourteenth International Conference on Learning Representations (ICLR 2026), 2026 · 31 pages, 2 figures

arxiv created 2026/07/30 · arxiv updated 2026/07/31

Abstract

As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity as context length n increases, causing tokens to cluster excessively, a phenomenon known as rank-collapse. While attention scaling effectively addresses this deficiency by rescaling attention scores with a polylogarithmic factor βn, theoretical justification for this approach remains lacking. We analyze a simplified yet tractable model that magnifies the effect of attention scaling. In this model, attention exhibits a phase transition governed by the scaling factor βn: insufficient scaling collapses all tokens to a single direction, while excessive scaling reduces attention to identity, thereby eliminating meaningful interactions between tokens. Our main result identifies the critical scaling βn \asymp log n and provides a rigorous justification for attention scaling in YaRN and Qwen, clarifying why logarithmic scaling maintains sparse, content-adaptive attention at large context lengths.

Citations

Cited by

Related