vix.ing · top · new · best · stats · spec

A Low-Resolution Image is Worth 1x1 Words: Enabling Fine Image Super-Resolution with Transformers and TaylorShift

2024/11/15 by Samudrala Nagaraju, Brian B. Moser, Nagaraju, Sanath Budakegowdanadoddi +9 · 1 citation
Computer Science · #Advanced Image Processing Techniques #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM)

paper · pdf · doi:10.48550/arxiv.2411.10231

openalex publication_date 2024/11/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Transformer-based architectures have recently advanced the image reconstruction quality of super-resolution (SR) models. Yet, their scalability remains limited by quadratic attention costs and coarse patch embeddings that weaken pixel-level fidelity. We propose TaylorIR, a plug-and-play framework that enforces 1x1 patch embeddings for true pixel-wise reasoning and replaces conventional self-attention with TaylorShift, a Taylor-series-based attention mechanism enabling full token interactions with near-linear complexity. Across multiple SR benchmarks, TaylorIR delivers state-of-the-art performance while reducing memory consumption by up to 60%, effectively bridging the gap between fine-grained detail restoration and efficient transformer scaling.

Cited by

Related