Reasoning Beyond Words ? Exploring framework for hidden state reasoning
2024/12/09 by Shibo Hao, Hao, Shibo, Sainbayar Sukhbaatar +11 · 40 voices · 177 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2412.06769
Abstract
Large language models (LLMs) are typically constrained to reason in the language space, where they express the reasoning process through a chain-of-thought (CoT) to solve complex problems. However, the language space may not always be optimal for reasoning. Most word tokens primarily ensure textual coherence and are not essential for reasoning, while some critical tokens require complex planning and pose challenges to LLMs. To explore the potential of reasoning beyond language, we introduce a new paradigm called Coconut (Chain of Continuous Thought). Coconut utilizes the last hidden state of the LLM as a representation of the reasoning state, termed "continuous thought." Instead of decoding this state into words, we feed it back to the model as the next input embedding directly in the continuous space. This latent reasoning paradigm enables an advanced reasoning pattern, where continuous thoughts can encode multiple alternative next steps, allowing the model to perform a breadth-first search (BFS) rather than committing prematurely to a single deterministic path as in CoT. Coconut outperforms CoT on logical reasoning tasks that require substantial search during planning and achieves a better trade-off between accuracy and efficiency.
Cited by
- J-CoT: Chain-of-Thought in J-Space
- Pretraining Recurrent Networks without Recurrence
- SHIFT: Self-reconstruction Harnesses Implicit Fine-grained Thinking for Retrieval
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
- SLPO: Scaling Latent Reasoning via a Surrogate Policy
- Memoir: Should a Model Write to Its Memory While It Thinks?
- HiCI: Hierarchical Construction-Integration for Long-Context Attention
- LatentMT: Machine Translation with Latent Reasoning
- Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary
- MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
- Constrained Path Reasoning: Measuring When Committed Stages Earn Their Cost
- T2MLR: Transformer with Temporal Middle-Layer Recurrence
- Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts
- MUX: Continuous Reasoning via Multiplexed Tokens
- Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought
- Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Hierarchical Reasoning Model
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought
- Reasoning Models Can Be Effective Without Thinking
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- Efficient Reasoning with Hidden Thinking
- Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Machine Learning Decoding of Circuit-Level Noise for Bivariate Bicycle Codes
- DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning
- The Case Against Generation for Retrieval: Discriminative Language Models as Effective Retrievers
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- Not All LLM Reasoning is Visible in the Chain-of-Thought
- Latent Implicit Visual Reasoning
- Observer, Not Player: Simulating Theory of Mind in LLMs through Game Observation
- JEPA-Reasoner: Decoupling Latent Reasoning from Token Generation
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
- From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition
- State over Tokens: Characterizing the Role of Reasoning Tokens
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- Latent Chain-of-Thought World Modeling for End-to-End Driving
- Mull-Tokens: Modality-Agnostic Latent Thinking
- Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
- Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
- Generative Recursive Reasoning
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems
- Unsupervised decoding of encoded reasoning using language model interpretability
- Lightweight Latent Reasoning for Narrative Tasks
- Difficulties with Evaluating a Deception Detector for AIs
- Learning When to Stop: Adaptive Latent Reasoning via Reinforcement Learning
- Reinforcement Learning for Latent-Space Thinking in LLMs
- Latent Collaboration in Multi-Agent Systems
- CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- Incorporating Self-Rewriting into Large Language Model Reasoning Reinforcement
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback
- SpiralThinker: Latent Reasoning through an Iterative Process with Text-Latent Interleaving
- Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
- The Future of Generative AI in Software Engineering: A Vision from Industry and Academia in the European GENIUS Project
- SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
- C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning
- Think Consistently, Reason Efficiently: Energy-Based Calibration for Implicit Chain-of-Thought
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- Beyond Single Embeddings: Capturing Diverse Targets with Multi-Query Retrieval
- EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities
- Penelope: Localized Latent Recurrence for Efficient Structured Reasoning
- Cache Merging as a Convergent Replicated State for Multi-Agent Latent Reasoning
- From Interface to Inference: Eliciting Any-Order Inference from Any-Order Models
- Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM
- HRM-Text: Efficient Pretraining Beyond Scaling
- Can Aha Moments Be Fake? Identifying True and Decorative Thinking Steps in Chain-of-Thought
- SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
- Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
- Parallel Loop Transformer for Efficient Test-Time Computation Scaling
- Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
- A Pragmatic Way to Measure Chain-of-Thought Monitorability
- HRM-Agent: Training a recurrent reasoning model in dynamic environments using reinforcement learning
- Context-level Language Modeling by Learning Predictive Context Embeddings
- A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring
- EffiReasonTrans: RL-Optimized Reasoning for Code Translation
- ActivationReasoning: Logical Reasoning in Latent Activation Spaces
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
- Soft-Masked Diffusion Language Models
- Planner and Executor: Collaboration between Discrete Diffusion And Autoregressive Models in Reasoning
- LLM Latent Reasoning as Chain of Superposition
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
- Towards Inference-time Scaling for Continuous Space Reasoning
- Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production
- Latent Refinement Decoding: Enhancing Diffusion-Based Language Models by Refining Belief States
- DND: Boosting Large Language Models with Dynamic Nested Depth
- Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning
- One Token Embedding Is Enough to Deadlock Your Large Reasoning Model
- Concise Reasoning in the Lens of Lagrangian Optimization
- Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers
- The Geometry of Reasoning: Flowing Logics in Representation Space
- Agentic Systems in Radiology: Design, Applications, Evaluation, and Challenges
- ReFIne: A Framework for Trustworthy Large Reasoning Models with Reliability, Faithfulness, and Interpretability
- Parallel Test-Time Scaling for Latent Reasoning Models
- R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
- MeSH: Memory-as-State-Highways for Recursive Transformers
- Can Speech LLMs Think while Listening?
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
- Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
- Efficient numeracy in language models through single-token number embeddings
- KaVa: Latent Reasoning via Compressed KV-Cache Distillation
- Incoherence in Goal-Conditioned Autoregressive Models
- MixReasoning: Switching Modes to Think
- SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
- KEEP: Integrating Medical Ontologies with Clinical Data for Robust Code Embeddings
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
- The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
- What Drives Compositional Generalization in Visual Generative Models?
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
- Typed Chain-of-Thought: A Curry-Howard Framework for Verifying LLM Reasoning
- Exploring System 1 and 2 communication for latent reasoning in LLMs
- Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- Hierarchical Reasoning Models: Perspectives and Misconceptions
- LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
- MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
- Latent Visual Reasoning
- Learning to Ponder: Adaptive Reasoning in Latent Space
- Deep Thinking by Markov Chain of Continuous Thoughts
- Alternatives To Next Token Prediction In Text Generation -- A Survey
- Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
- Efficient Turing Machine Simulation with Transformers
- Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought
- Two-Scale Latent Dynamics for Recurrent-Depth Transformers
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- A model of errors in transformers
- MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
- R-Capsule: Compressing High-Level Plans for Efficient Large Language Model Reasoning
- Learning to Reason with Mixture of Tokens
- A Formal Comparison Between Chain-of-Thought and Latent Thought
- SIM-CoT: Supervised Implicit Chain-of-Thought
- Interpretable Representation via LLM-Driven Generative Disentanglement for Local-Life Service Recommendation
- Hierarchical Latent Reasoning for LLM-based Recommendation
- OPLD: On-Policy Latent Distillation for Multimodal Reasoning
- Can AI Follow In Einstein's Footsteps?
- SuperThoughts: Reasoning Tokens in Superposition
- LASAR: Latent Adaptive Semantic Aligned Reasoning for Generative Recommendation
- The Topological Trouble With Transformers
- Soft Tokens, Hard Truths
- OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
- Stop Spinning Wheels: Mitigating LLM Overthinking via Mining Patterns for Early Reasoning Exit
- VOX-KRIKRI: Unifying Speech and Language through Continuous Fusion
- Meta-R1: Empowering Large Reasoning Models with Metacognition
- Towards mitigating information leakage when evaluating safety monitors
- LTA-thinker: Latent Thought-Augmented Training Framework for Large Language Models on Complex Reasoning
- Investigating Language Model Capabilities to Represent and Process Formal Knowledge: A Preliminary Study to Assist Ontology Engineering
- Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
- Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
- From Long to Short: LLMs Excel at Trimming Own Reasoning Chains
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
- Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics
- Reliable Weak-to-Strong Monitoring of LLM Agents
- Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
- Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
- Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
- Pruning the Unsurprising: Efficient LLM Reasoning via First-Token Surprisal
- Navigating Through Paper Flood: Advancing LLM-based Paper Evaluation through Domain-Aware Retrieval and Latent Reasoning
- LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking
- Compressing Chain-of-Thought in LLMs via Step Entropy
- SynAdapt: Learning Adaptive Reasoning in Large Language Models via Synthetic Continuous Chain-of-Thought
- Integrating clinical reasoning into large language model-based diagnosis through etiology-aware attention steering
- Explainability Through Systematicity: The Hard Systematicity Challenge for Artificial Intelligence
- Feedback neural network [wikipedia]
Discussions
- Training LLMs to Reason in a Continuous Latent Space [hn, 283 points, 114 comments]
- Sometimes our anthropocentric assumptions about how intelligence "should" work (like using language for reasoning) may be holding AI back. Letting AI reason in its own native "language" in latent spac [bsky, 94 points, 5 comments]
- Training Large Language Models to Reason in a Continuous Latent Space Introduces a new paradigm for LLM reasoning called Chain of Continuous Thought (COCONUT) Directly feed the last hidden state (a [bsky, 54 points, 3 comments]
- uh oh. they're working on model architectures that reason and plan directly in latent space instead of using word-based chain of thought. Coconut and JEPA for example. we're neck-deep in unsettling co [bsky, 14 points, 2 comments]
- btw this was meant to try skipping tokenisation altogether, but it didn't seem to take off: arxiv.org/abs/2412.06769 [bsky, 10 points, 1 comments]
- Training Large Language Models to Reason in a Continuous Latent Space https://arxiv.org/abs/2412.06769v2 (attached screenshot from Andrew Ng’s newsletter) #AI #reasoning [bsky, 7 points, 0 comments]
- I always thought that reasoning does not require language. Well, this seems to be supported by neuroscience, see screenshot from arxiv.org/pdf/2412.06769 [bsky, 5 points, 0 comments]
- Pretty interesting paper by Meta & UC San Diego: arxiv.org/abs/2412.067... unlike humans, which don't activate the language network for many reasoning tasks, LLM are forced to reason in "language spac [bsky, 3 points, 1 comments]
- i have this paper in my queue. they basically just cut out the final layer that projects into logits and directly cycle the hidden layer output to the front of the LLM seems like what you’re looking [bsky, 3 points, 1 comments]
- A la arxiv.org/abs/2412.06769 [bsky, 2 points, 1 comments]
- Very cool paper on the internal dynamics of reasoning in LMs. The approach (Chain of Continuous Thought) lets models reason in continuous latent space rather than being constrained to generating speci [bsky, 2 points, 1 comments]
- Ce papier est fascinant et montre qu'une IA qu'on laisse définir des nouveaux concepts hors langage raisonne mieux. Ca fait sens : sur des problèmes complexes de math, je ne raisonne pas du tout en fr [bsky, 2 points, 0 comments]
- Your point about alien language reminded me of this paper: arxiv.org/abs/2412.06769 The LLMs are still trained in human languages but this 'continuous chain of thought approach' keeps the reasoning [bsky, 2 points, 0 comments]
- My attention was drawn recently to this paper. Some ML folks not only identified the fundamental design issue with LLMs which I've been going on about for a while, but it seems they also identified a [bsky, 1 points, 1 comments]
- There is some extra information you could preserve by saving the full-dimensional output at each position and using that as input instead of a token- that’s what arxiv.org/abs/2412.06769 is. It’s uncl [bsky, 1 points, 1 comments]
- I'm aware of these two recent papers implementing reasoning in latent space: - arxiv.org/abs/2412.06769 - arxiv.org/abs/2412.17747 [bsky, 1 points, 1 comments]
- Models can probably get better performance by reasoning in embedding space rather than in tokens (arxiv.org/pdf/2412.06769), I think we'll probably see large-scale models that reason entirely in laten [bsky, 1 points, 2 comments]
- Using the last state of the model as input without actually generating a token always seemed like an obvious idea. I guess it just took a long time to finish the paper. Great to see some results now. [bsky, 1 points, 0 comments]
- Here they utilize the last hidden state of the LLM as a representation of the reasoning state (termed "continuous thought"). Rather than decoding this into a word token, they feed it back to the LLM a [bsky, 1 points, 0 comments]
- The next paper I saw was on continuous chain of thought, creating new latent thoughts that are much more expressive and allow the model to compress its thinking by an OOM. arxiv.org/abs/2412.06769 [bsky, 1 points, 1 comments]
- Training Large Language Models to Reason in a Continuous Latent Space [pdf] [hn, 1 points, 0 comments]
- Training LLMs to Reason in a Continuous Latent Space https://arxiv.org/abs/2412.06769 (https://news.ycombinator.com/item?id=42378335) [bsky, 0 points, 0 comments]
- I think I only understand like 50% of this but it's v cool arxiv.org/abs/2412.06769 - the general idea of feeding latent cognition back into the inference instead of having every input being a previou [bsky, 0 points, 1 comments]
- Training LLMs to Reason in a Continuous Latent Space [bsky, 0 points, 0 comments]
- Meta and UC San Diego's Coconut framework enhances LLM reasoning by utilizing continuous latent space for improved planning and efficiency, allows reasoning without being limited by words. arxiv.org/a [bsky, 0 points, 0 comments]
- Interesting paper from Meta shows how letting AI models reason directly in neural space, rather than through tokens, leads to more efficient and flexible problem-solving arxiv.org/pdf/2412.06769 [bsky, 0 points, 0 comments]
- I don't think that's the current prevalent narrative in many circles. So much focus now is on reasoning models which is quite a different thing from a single pass through an LLM. And the new paper on [bsky, 0 points, 1 comments]
- Paper from Meta that describes “continuous thinking in latent space” in a way that can’t be done with GPT and chain of thought (CoT) reasoning. arxiv.org/pdf/2412.06769 *Worse* than Chain-of-thought [bsky, 0 points, 0 comments]
- Training LLMs to Reason in a Continuous Latent Space Research exploring advanced reasoning capabilities for large language models, potentially improving AI reasoning and problem-solving Read here [bsky, 0 points, 0 comments]
- Training LLMs to Reason in a Continuous Latent Space (arxiv.org) Main Link | Discussion [bsky, 0 points, 0 comments]
- Training Large Language Models to Reason in a Continuous Latent Space "We utilize the last hidden state of the LLM as a representation of the reasoning state...we feed it back to the LLM as the subseq [bsky, 0 points, 0 comments]
- Would love to see more work on reasoning in the concept space for LLM, like in this paper: arxiv.org/abs/2412.06769 [bsky, 0 points, 0 comments]
- Training LLMs to Reason in a Continuous Latent Space https://arxiv.org/abs/2412.06769 (https://news.ycombinator.com/item?id=42378335) [bsky, 0 points, 0 comments]
- How about 'Breadth-First Latent Reasoning'? Additionally, this represents a significant improvement in LLM reasoning. arxiv.org/abs/2412.06769 [bsky, 0 points, 1 comments]
- No More Words: Does reasoning require language? A new paper suggests that for certain kinds of problems AI reasoning models are better off leaving words behind. arxiv.org/abs/2412.06769 [bsky, 0 points, 1 comments]
- Training LLMs to Reason in a Continuous Latent Space https://arxiv.org/abs/2412.06769 [comments] [116 points] [bsky, 0 points, 0 comments]
- Training LLMs to Reason in a Continuous Latent Space https://arxiv.org/abs/2412.06769 arxiv.org [bsky, 0 points, 0 comments]
- Here is a “demo” that, given a tradeoff between AI transparency (English-language chain-of-thought) and AI capability (inscrutable chain-of-thought but the results are better), many people will choose [bsky, 0 points, 1 comments]
- What really grabbed me is how COCONUT mimics a "thought process" through these continuous latent trajectories. It’s like the AI is creating its own mental map to solve complex tasks! The planning par [bsky, 0 points, 2 comments]
- Training LLMs to Reason in a Continuous Latent Space https://arxiv.org/abs/2412.06769 https://news.ycombinator.com/item?id=42378335 [bsky, 0 points, 0 comments]
Related