vix.ing · top · new · best · stats · spec

Transformer Working Memory Enables Regular Language Reasoning and Natural Language Length Extrapolation

2023/05/05 by Ta-Chung Chi, Ting-Han Fan, Chi, Ta-Chung +5 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2305.03796

openalex publication_date 2023/05/05 · openalex created_date 2023/05/10 · openalex updated_date 2026/07/28

Abstract

Unlike recurrent models, conventional wisdom has it that Transformers cannot perfectly model regular languages. Inspired by the notion of working memory, we propose a new Transformer variant named RegularGPT. With its novel combination of Weight-Sharing, Adaptive-Depth, and Sliding-Dilated-Attention, RegularGPT constructs working memory along the depth dimension, thereby enabling efficient and successful modeling of regular languages such as PARITY. We further test RegularGPT on the task of natural language length extrapolation and surprisingly find that it rediscovers the local windowed attention effect deemed necessary in prior work for length extrapolation.

Cited by

Related