Reasoning with Latent Thoughts: On the Power of Looped Transformers
2025/02/24 by Nikunj Saunshi, Nishanth Dikkala, Saunshi, Nikunj +7 · 73 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Semantic Web and Ontologies
paper · pdf · doi:10.48550/arxiv.2502.17416
openalex publication_date 2025/02/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large language models have shown remarkable reasoning abilities and scaling laws suggest that large parameter count, especially along the depth axis, is the primary driver. In this work, we make a stronger claim -- many reasoning problems require a large depth but not necessarily many parameters. This unlocks a novel application of looped models for reasoning. Firstly, we show that for many synthetic reasoning problems like addition, p-hop induction, and math problems, a k-layer transformer looped L times nearly matches the performance of a kL-layer non-looped model, and is significantly better than a k-layer model. This is further corroborated by theoretical results showing that many such reasoning problems can be solved via iterative algorithms, and thus, can be solved effectively using looped models with nearly optimal depth. Perhaps surprisingly, these benefits also translate to practical settings of language modeling -- on many downstream reasoning tasks, a language model with k-layers looped L times can be competitive to, if not better than, a kL-layer language model. In fact, our empirical analysis reveals an intriguing phenomenon: looped and non-looped models exhibit scaling behavior that depends on their effective depth, akin to the inference-time scaling of chain-of-thought (CoT) reasoning. We further elucidate the connection to CoT reasoning by proving that looped models implicitly generate latent thoughts and can simulate T steps of CoT with T loops. Inspired by these findings, we also present an interesting dichotomy between reasoning and memorization, and design a looping-based regularization that is effective on both fronts.
Cited by
- DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning
- How Transformers Learn to Plan via Multi-Token Prediction
- VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
- Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
- ViT3: Unlocking Test-Time Training in Vision
- SpiralThinker: Latent Reasoning through an Iterative Process with Text-Latent Interleaving
- Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- LaRe: Latent Refocusing for Multimodal Reasoning
- HRM-Text: Efficient Pretraining Beyond Scaling
- Parallel Loop Transformer for Efficient Test-Time Computation Scaling
- On the Reasoning Abilities of Masked Diffusion Language Models
- Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
- Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
- DND: Boosting Large Language Models with Dynamic Nested Depth
- Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers
- Upfront Chain-of-Thought: A Cooperative Framework for Chain-of-Thought Compression
- MeSH: Memory-as-State-Highways for Recursive Transformers
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Models
- Fundamental Limits of Crystalline Equivariant Graph Neural Networks: A Circuit Complexity Perspective
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- The Thinking Spectrum: An Empirical Study of Tunable Reasoning in LLMs through Model Merging
- A Formal Comparison Between Chain-of-Thought and Latent Thought
- SIM-CoT: Supervised Implicit Chain-of-Thought
- Looped Transformers with Source-Centered State Evolution
- LASAR: Latent Adaptive Semantic Aligned Reasoning for Generative Recommendation
- The Topological Trouble With Transformers
- Soft Tokens, Hard Truths
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Towards High-Order Mean Flow Generative Models: Feasibility, Expressivity, and Provably Efficient Criteria
- ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding
- Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
- Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey
- A Survey on Latent Reasoning
- Think How to Think: Mitigating Overthinking with Autonomous Difficulty Cognition in Large Reasoning Models
- Energy-Based Transformers are Scalable Learners and Thinkers
- Efficient Post-Training Refinement of Latent Reasoning in Large Language Models
- Working Memory as Programmable Fast Weight Computation
- Parallel Continuous Chain-of-Thought with Jacobi Iteration
- Beyond Parameters: Exploring Virtual Logic Depth for Scaling Laws
- BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
- Representation Consistency for Accurate and Coherent LLM Answer Aggregation
- Scalable Chain of Thoughts via Elastic Reasoning
- Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning
- OThink-R1: Intrinsic Fast/Slow Thinking Mode Switching for Over-Reasoning Mitigation
- SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-Thought
- Continuous Chain of Thought Enables Parallel Exploration and Reasoning
- On Learning Verifiers and Implications to Chain-of-Thought Reasoning
- Pretraining Language Models to Ponder in Continuous Space
- Recurrent Self-Attention Dynamics: An Energy-Agnostic Perspective from Jacobians
- To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
- Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
- TextNCA: Neural Cellular Automata for Language Modeling via Hierarchical Local Attention
- Multilingual Test-Time Scaling via Initial Thought Transfer
- Thinkless: LLM Learns When to Think
- Accelerating Chain-of-Thought Reasoning: When Goal-Gradient Importance Meets Dynamic Skipping
- Adaptive Loops and Memory in Transformers: Think Harder or Know More?
- Looped World Models
- Universal Transformers Need Memory: Depth-State Trade-offs in Adaptive Recursive Reasoning
- Training-Free Looped Transformers
- Solve the Loop: Attractor Models for Language and Reasoning
- Optimal Order of Multi-Agent and General Many-Body Systems
- Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
- Interleaved Head Attention
- Efficient Reasoning for LLMs through Speculative Chain-of-Thought
- LoopMTP: A looped transformer guided by latent multi-token prediction
- Training Crossroads for Recurrent Vision Transformers: Recurrence, Neural ODEs, and Deep Supervision
- Dynamic Early Exit in Reasoning Models
- Iterate or Widen? When Test-Time Refinement Helps LiDAR Scene Completion: A Controlled Study of Evidence Geometry, Training Coverage, and Compute
- Efficient Reasoning Models: A Survey
Related