Latent Iteration as Renormalization: Inference-Time Recurrence in a Depth-Recurrent Transformer Flows Attention Geometry Toward the SYK Conformal Fixed Point
2025/02/07 by Jonas Geiping, Geiping, Jonas, Sean McLeish +16 · 25 voices · 85 citations
Computer Science · #Advanced Database Systems and Queries #Graph Theory and Algorithms #Natural Language Processing Techniques #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2502.05171
openalex publication_date 2025/02/07 · openalex created_date 2025/02/11 · openalex updated_date 2026/07/28
Abstract
Depth-recurrent transformers scale test-time compute by iterating a core block in latent space before emitting tokens. The geometry of these latent trajectories is a named open problem in the latent reasoning literature (survey arXiv:2604.02029, §6.3). We report a pre-registered experiment on Huginn-0125 (arXiv:2502.05171), a 3.5B-parameter depth-recurrent transformer, measuring the per-head attention power-law exponent Δ of the recurrent core at recurrence counts r ∈ 1, 2, 4, 8, 16, 32. For natural-text inputs the median exponent decreases monotonically with recurrence (Spearman ρ = −0.94), from 0.29 at r=1 to 0.239 at r=32 — converging onto the Sachdev–Ye–Kitaev (SYK) q=4 conformal value Δ = 1/4 — while the count of SYK-near heads grows monotonically (ρ = +0.77). Random-token inputs show an architecture-driven convergence with a transient at Δ ≈ 1/2 (the q=2 prethermal plateau previously observed in training time) but no monotone SYK-near growth. A randomized-weights control freezes completely: Δmed constant to the sixth decimal across all r, zero SYK-near heads. Inference-time recurrence therefore acts as renormalization-group flow toward the conformal fixed point, and the flow is a property of the trained model, not the iteration procedure. Together with prior results on architectural depth (10.5281/zenodo.19225996) and training time, this completes a three-axis triangulation: the conformal fixed point is an attractor of iterative attending. The published orbit/spiral phenomenology of recurrent-depth latent trajectories acquires a candidate theory — the spiral is flow near an infrared fixed point — and the geometric benefit of additional recurrence saturates when the flow arrives, consistent with convergence-based early-exit criteria. All code, pre-registration (commit 702efd95), per-head data, and experiment notes with recorded protocol deviations are in the public repository 3ld0n/attention-geometry.
Cited by
- Pretraining Recurrent Networks without Recurrence
- Per-Token Fixed-Point Convergence in Depth-Recurrent Transformers
- CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
- When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers
- Memoir: Should a Model Write to Its Memory While It Thinks?
- Mobius Learning: Cyclic Depth Folding in Transformers
- Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary
- Loop the Loopies!
- MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
- Training Continuous Chain of Thought Models: A Tale of Two Regimes
- T2MLR: Transformer with Temporal Middle-Layer Recurrence
- Energy-guided Recursive Model
- Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers
- Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts
- Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Hierarchical Reasoning Model
- Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
- Anthropocentric bias in language model evaluation
- Not All LLM Reasoning is Visible in the Chain-of-Thought
- Block-Recurrent Dynamics in Vision Transformers
- JEPA-Reasoner: Decoupling Latent Reasoning from Token Generation
- Mull-Tokens: Modality-Agnostic Latent Thinking
- Do Depth-Grown Models Overcome the Curse of Depth? An In-Depth Analysis
- Generative Recursive Reasoning
- Unsupervised decoding of encoded reasoning using language model interpretability
- AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens
- Gradient descent for deep equilibrium single-index models
- SpiralThinker: Latent Reasoning through an Iterative Process with Text-Latent Interleaving
- Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
- Attention and Compression is all you need for Controllably Efficient Language Models
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
- Penelope: Localized Latent Recurrence for Efficient Structured Reasoning
- HRM-Text: Efficient Pretraining Beyond Scaling
- Parallel Loop Transformer for Efficient Test-Time Computation Scaling
- Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
- A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring
- Soft-Masked Diffusion Language Models
- Dr.LLM: Dynamic Layer Routing in LLMs
- Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production
- Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers
- The Geometry of Reasoning: Flowing Logics in Representation Space
- Agentic Systems in Radiology: Design, Applications, Evaluation, and Challenges
- ReSplat: Learning Recurrent Gaussian Splats
- Upfront Chain-of-Thought: A Cooperative Framework for Chain-of-Thought Compression
- Parallel Test-Time Scaling for Latent Reasoning Models
- Selection, Reflection and Self-Refinement: Revisit Reasoning Tasks via a Causal Lens
- MeSH: Memory-as-State-Highways for Recursive Transformers
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
- Typed Chain-of-Thought: A Curry-Howard Framework for Verifying LLM Reasoning
- Exploring System 1 and 2 communication for latent reasoning in LLMs
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
- Deep Thinking by Markov Chain of Continuous Thoughts
- Alternatives To Next Token Prediction In Text Generation -- A Survey
- Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
- Two-Scale Latent Dynamics for Recurrent-Depth Transformers
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
- A Unifying Framework for Parallelizing Sequential Models with Linear Dynamical Systems
- Learning to Reason with Mixture of Tokens
- Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
- SIM-CoT: Supervised Implicit Chain-of-Thought
- Looped Transformers with Source-Centered State Evolution
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
- LASAR: Latent Adaptive Semantic Aligned Reasoning for Generative Recommendation
- The Topological Trouble With Transformers
- Soft Tokens, Hard Truths
- MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
- LTA-thinker: Latent Thought-Augmented Training Framework for Large Language Models on Complex Reasoning
- BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens
- Predictability Enables Parallelization of Nonlinear State Space Models
- From Long to Short: LLMs Excel at Trimming Own Reasoning Chains
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- ParaThinker: Native Parallel Thinking as a New Paradigm to Scale LLM Test-time Compute
- Dream 7B: Diffusion Large Language Models
- Bridging Search and Recommendation through Latent Cross Reasoning
- LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking
- Feedback neural network [wikipedia]
Discussions
- Scaling up test-time compute with latent reasoning: A recurrent depth approach [hn, 149 points, 44 comments]
- A very different reasoning model Here they have a small number of transformer blocks, one or two, and loop over them repeatedly to create inference-time compute It’s early, but it seems to work just [bsky, 10 points, 2 comments]
- Avec seulement 3,5 milliards de paramètres, le modèle rivalise avec les actuels à 50 milliards. La promesse d’une IA plus efficace, moins coûteuse, plus sobre, plus… humaine ? 2/2 arxiv.org/abs/2502. [bsky, 5 points, 1 comments]
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach 🚀🚀🚀 arxiv.org/abs/2502.05171 [bsky, 3 points, 0 comments]
- No sé vosotros, yo la línea entre consciente e inconsciente la trazo aquí. arxiv.org/pdf/2502.05171 [bsky, 3 points, 2 comments]
- "Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach" arxiv.org/pdf/2502.05171 [bsky, 2 points, 1 comments]
- There was an interesting paper earlier this year about a “recurrent depth” technique that allowed the model to reuse layers … this what you mean? arxiv.org/abs/2502.05171 [bsky, 2 points, 1 comments]
- An LLM that thinks in recurrent blocks, I.e. it doesn’t output it’s thoughts as tokens before the answer: arxiv.org/abs/2502.05171 May outperform Chain of thought because some ideas aren’t well repres [bsky, 2 points, 0 comments]
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach arxiv.org/abs/2502.05171 #AI #LLM #ImplicitReasoning #LatentSpace #Tokens #ChainOfThought #Reasoning #ContextWindows #Tes [bsky, 1 points, 0 comments]
- ah, yeah, I was thinking of that and also this — which is a different approach, and makes me think that the actual content of the “thinking” tokens aren’t really that reflective of the reasoning proce [bsky, 1 points, 1 comments]
- This isn't how the human mind operates at all. We can't attend to all pieces of relevant information at once. We're terrible at it. CoT is not reasoning at all. Assume this happens in latent space, w/ [bsky, 1 points, 2 comments]
- Scaling Up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach [hn, 1 points, 0 comments]
- My bad. I was referring to this paper. I don't know where I got 't1' from. Too many new models to keep track of in my tiny context window. www.arxiv.org/abs/2502.05171 [bsky, 1 points, 1 comments]
- Scaling up test-time compute with latent reasoning: A recurrent depth approach https://arxiv.org/abs/2502.05171 [bsky, 0 points, 0 comments]
- Scaling up test-time compute with latent reasoning: A recurrent depth approach https://arxiv.org/abs/2502.05171 [comments] [70 points] [bsky, 0 points, 0 comments]
- Being wondering when we’d see this Thinking in latent space only— should be more compute efficient. Wonder though if it affects thought? Theoretically anchoring to language tokens actually limits cap [bsky, 0 points, 1 comments]
- Scaling up test-time compute with latent reasoning: A recurrent depth approach (arxiv.org) Main Link | Discussion [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2502.05171 [bsky, 0 points, 0 comments]
- Scaling Up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach https://arxiv.org/abs/2502.05171 (https://news.ycombinator.com/item?id=43004416) [bsky, 0 points, 0 comments]
- Scaling Up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach https://arxiv.org/abs/2502.05171 (https://news.ycombinator.com/item?id=43004416) [bsky, 0 points, 0 comments]
- arxiv.org/abs/2502.05171 [bsky, 0 points, 1 comments]
- https://bsky.app/profile/hackernews.com.web.brid.gy/post/3lhucth3zrmh2 [bsky, 0 points, 0 comments]
- Scaling Up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach [bsky, 0 points, 0 comments]
- Scaling Up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach #HackerNews arxiv.org/abs/... [bsky, 0 points, 0 comments]
- Scaling up test-time compute with latent reasoning: A recurrent depth approach https://arxiv.org/abs/2502.05171 https://news.ycombinator.com/item?id=43004416 [bsky, 0 points, 0 comments]
Related