2024/10/02 by Kevin Xu, Xu, Kevin, Issei Sato +1 · 3 citations
Computer Science · Engineering · #Advanced Memory and Neural Computing #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications #Neural Networks and Reservoir Computing
paper · pdf · doi:10.48550/arxiv.2410.01405
openalex publication_date 2024/10/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Looped Transformers provide advantages in parameter efficiency, computational capabilities, and generalization for reasoning tasks. However, their expressive power regarding function approximation remains underexplored. In this paper, we establish the approximation rate of Looped Transformers by defining the modulus of continuity for sequence-to-sequence functions. This reveals a limitation specific to the looped architecture. That is, the analysis prompts the incorporation of scaling parameters for each loop, conditioned on timestep encoding. Experiments validate the theoretical results, showing that increasing the number of loops enhances performance, with further gains achieved through the timestep encoding.