2021/08/03 by Lingchuan Meng, Meng, Lingchuan · 2 citations
Computer Science · Engineering · Mathematics · #Advanced Memory and Neural Computing #Advanced Neural Network Applications #Armour #Artificial intelligence #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Domain Adaptation and Few-Shot Learning #Electrical engineering #Engineering #FOS: Computer and information sciences #Geometry #Materials science #Mathematics #Nanotechnology #Quadratic equation #Transformer #Voltage #cs.CV
paper · pdf · doi:10.48550/arxiv.2108.01778
published in arXiv (Cornell University) (Cornell University)
arxiv created 2021/08/03 · openalex publication_date 2021/08/03 · arxiv updated 2021/08/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Attention-based transformer networks have demonstrated promising potential as their applications extend from natural language processing to vision. However, despite the recent improvements, such as sub-quadratic attention approximation and various training enhancements, the compact vision transformers to date using the regular attention still fall short in comparison with its convnet counterparts, in terms of accuracy, model size, and throughput. This paper introduces a compact self-attention mechanism that is fundamental and highly generalizable. The proposed method reduces redundancy and improves efficiency on top of the existing attention optimizations. We show its drop-in applicability for both the regular attention mechanism and some most recent variants in vision transformers. As a result, we produced smaller and faster models with the same or better accuracies.