2025/10/02 by Haochen You, You, Haochen, Baojing Liu +1
Computer Science · #AI-based Problem Solving and Planning #Bounded function #Computation and Language (cs.CL) #Encoder #FLOPS #FOS: Computer and information sciences #Generalization #Inference #Iterative refinement #Logic, Reasoning, and Knowledge #Networking and Internet Architecture (cs.NI) #Scalability #Semantic Web and Ontologies #Transformer
paper · pdf · doi:10.48550/arxiv.2510.01585
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/10/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
While Transformer architectures have demonstrated impressive scalability across domains, they continue to face challenges in long-context reasoning, computational efficiency, and structural generalization - largely due to rigid layer stacking, dense attention, and reliance on positional encodings. We present ReSSFormer, a Recursive Sparse Structured Transformer that integrates three complementary innovations: Recurrent Reasoning & Memory Unit (R2MU) for iterative reasoning with bounded depth, Adaptive Sparse Attention Module (ASAM) for efficient and focused context selection, and Self-Organizing Encoder Structure (SOES) for position-free structure induction. ReSSFormer replaces conventional depth stacking with recurrent inference, substitutes full attention with token- and expert-level sparsity, and models latent token topology directly from content. Across language modeling, multi-hop QA, and structure-sensitive tasks, ReSSFormer consistently outperforms strong baselines under comparable FLOPs and parameter budgets, highlighting its scalability, efficiency, and structural flexibility.