2025/02/04 by Chaofan Lin, Jiaming Tang, Lin, Chaofan +15 · 1 voice · 14 citations
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #Intelligent Tutoring Systems and Adaptive Learning #Online Learning and Analytics #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2502.02770
arxiv published 2025/02/04 · arxiv updated 2025/11/04
Leveraging attention sparsity to accelerate long-context large language models (LLMs) has been a hot research topic. However, current algorithms such as sparse attention or key-value (KV) cache compression tend to use a fixed budget, which presents a significant challenge during deployment because it fails to account for the dynamic nature of real-world scenarios, where the optimal balance between accuracy and efficiency can vary greatly. In this paper, we find that borrowing top-p sampling (nucleus sampling) to sparse attention can surprisingly achieve adaptive budgeting. Based on this, we propose Twilight, a framework to bring adaptive sparsity to any existing sparse attention algorithm without sacrificing their accuracy. Empirical results show that Twilight can adaptively prune at most 98% of redundant tokens, leading to 15.4× acceleration in self-attention operations and 3.9× acceleration in end-to-end per token latency in long context LLM decoding.