Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience
2025/05/07 by Zinan Liu, Liu, Zinan, Haoran Li +15 · 2 citations
Computer Science · Neuroscience · Psychology · #Action Observation and Synchronization #Embodied and Extended Cognition #FOS: Biological sciences #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Neurons and Cognition (q-bio.NC)
paper · pdf · doi:10.48550/arxiv.2505.05515
openalex publication_date 2025/05/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Autonomous AI is no longer a hard-to-reach concept, it enables the agents to move beyond executing tasks to independently addressing complex problems, adapting to change while handling the uncertainty of the environment. However, what makes the agents truly autonomous? It is agentic reasoning, that is crucial for foundation models to develop symbolic logic, statistical correlations, or large-scale pattern recognition to process information, draw inferences, and make decisions. However, it remains unclear why and how existing agentic reasoning approaches work, in comparison to biological reasoning, which instead is deeply rooted in neural mechanisms involving hierarchical cognition, multimodal integration, and dynamic interactions. In this work, we propose a novel neuroscience-inspired framework for agentic reasoning. Grounded in three neuroscience-based definitions and supported by mathematical and biological foundations, we propose a unified framework modeling reasoning from perception to action, encompassing four core types, perceptual, dimensional, logical, and interactive, inspired by distinct functional roles observed in the human brain. We apply this framework to systematically classify and analyze existing AI reasoning methods, evaluating their theoretical foundations, computational designs, and practical limitations. We also explore its implications for building more generalizable, cognitively aligned agents in physical and virtual environments. Finally, building on our framework, we outline future directions and propose new neural-inspired reasoning methods, analogous to chain-of-thought prompting. By bridging cognitive neuroscience and AI, this work offers a theoretical foundation and practical roadmap for advancing agentic reasoning in intelligent systems. The associated project can be found at: https://github.com/BioRAILab/Awesome-Neuroscience-Agent-Reasoning .
Citations
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)
- Z1: Efficient Test-time Scaling with Code
- ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
- Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- MastermindEval: A Simple But Scalable Reasoning Benchmark
- Visual-RFT: Visual Reinforcement Fine-Tuning
- BIG-Bench Extra Hard
- MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning
- Chain of Draft: Thinking Faster by Writing Less
- From System 1 to System 2: A Survey of Reasoning Large Language Models
- End-to-End Chart Summarization via Visual Chain-of-Thought in Vision-Language Models
- Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models
- Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
- Reasoning with Reinforced Functional Token Tuning
- AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors
- LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
- SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
- Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding
- VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving
- Qwen2.5 Technical Report
- Meta-Reflection: A Feedback-Free Reflection Learning Framework
- SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World
- Interleaved-Modal Chain-of-Thought
- TopV-Nav: Unlocking the Top-View Spatial Reasoning Potential of MLLM for Zero-shot Object Navigation
- Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
- LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
- End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
- CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent Cooperation
- Vision-Language Models Can Self-Improve Reasoning via Reflection
- GPT-4o System Card
- RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
- Structured Spatial Reasoning with Open Vocabulary Object Detectors
- Temporal Reasoning Transfer from Text to Video
- Enhancing Temporal Sensitivity and Reasoning for Time-Sensitive Question Answering
- Language Grounded Multi-agent Reinforcement Learning with Human-interpretable Communication
- To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
- Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
- Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents
- ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning
- I Know About "Up"! Enhancing Spatial Reasoning in Visual Language Models Through 3D Reconstruction
- GRASP: A Grid-Based Benchmark for Evaluating Commonsense Spatial Reasoning
- Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
- MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs
- Transferable Tactile Transformers for Representation Learning Across Diverse Sensors and Tasks
- Abstraction-of-Thought Makes Language Models Better Reasoners
- LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
- ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search
- Touch100k: A Large-Scale Touch-Language-Vision Dataset for Touch-Centric Multimodal Representation
- Improve Mathematical Reasoning in Language Models by Automated Process Supervision
- SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
- Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism
- Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
- Octopi: Object Property Reasoning with Large Tactile-Language Models
- Self-playing Adversarial Language Game Enhances LLM Reasoning
- Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
- Vision-Language Model-based Physical Reasoning for Robot Liquid Perception
- Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
- HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning
- SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
- JSTR: Joint Spatio-Temporal Reasoning for Event-based Moving Object Detection
- Chain-of-Thought Reasoning Without Prompting
- GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering
- BAT: Learning to Reason about Spatial Sounds with Large Language Models
- HAZARD Challenge: Embodied Decision Making in Dynamically Changing Environments
- SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
- Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge
- Large Language Models Can Learn Temporal Reasoning
- Learning Audio Concepts from Counterfactual Natural Language
- Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
- Chain of Code: Reasoning with a Language Model-Augmented Code Emulator
- Visual cognition in multimodal large language models
- Improving Zero-shot Visual Question Answering via Large Language Models with Reasoning Question Prompts
- Large Language Models are Visual Reasoning Coordinators
- 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
- Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
- MenatQA: A New Dataset for Testing the Temporal Comprehension and Reasoning Abilities of Large Language Models
- TRAM: Benchmarking Temporal Reasoning for Large Language Models
- Enhancing Zero-Shot Chain-of-Thought Reasoning in Large Language Models through Logic
- Hypothesis Search: Inductive Reasoning with Language Models
- Temporal Inductive Path Neural Network for Temporal Knowledge Graph Reasoning
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
- LISA: Reasoning Segmentation via Large Language Model
- Discovering Spatio-Temporal Rationales for Video Question Answering
- RoCo: Dialectic Multi-Robot Collaboration with Large Language Models
- Building Cooperative Embodied Agents Modularly with Large Language Models
- Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models
- Inductive reasoning in humans and large language models
- Deductive Verification of Chain-of-Thought Reasoning
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
- Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD Examples
- Listen, Think, and Understand
- ReasonNet: End-to-End Driving with Temporal and Global Reasoning
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
- DERA: Enhancing Large Language Model Completions with Dialog-Enabled Resolving Agents
- TKN: Transformer-based Keypoint Prediction Network For Real-time Video Prediction
- GPT-4 Technical Report
- ViperGPT: Visual Inference via Python Execution for Reasoning
- Text-to-ECG: 12-Lead Electrocardiogram Synthesis conditioned on Clinical Text Reports
- LLaMA: Open and Efficient Foundation Language Models
- Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data
- Active Prompting with Chain-of-Thought for Large Language Models
- Large Language Models Are Reasoning Teachers
- Super-CLEVR: A Virtual Benchmark to Diagnose Domain Robustness in Visual Reasoning
- Temporal Knowledge Graph Reasoning with Historical Contrastive Learning
- Visual Programming: Compositional visual reasoning without training
- Zero-shot visual reasoning through probabilistic analogical mapping
- PromptCast: A New Prompt-based Learning Paradigm for Time Series Forecasting
- FOLIO: Natural Language Reasoning with First-Order Logic
- PEER: A Collaborative Language Model
- Can Brain Signals Reveal Inner Alignment with Human Languages?
- A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
- StreamingQA: A Benchmark for Adaptation to New Knowledge over Time in Question Answering Models
- GRIT: General Robust Image Task Benchmark
- PaLM: Scaling Language Modeling with Pathways
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- A Survey on Neural-symbolic Learning Systems
- A survey on neural-symbolic learning systems
- Training Verifiers to Solve Math Word Problems
- A Dataset for Answering Time-Sensitive Questions
- From LSAT: The Progress and Challenges of Complex Reasoning
- Time-Aware Language Models as Temporal Knowledge Bases
- AST: Audio Spectrogram Transformer
- Roses Are Red, Violets Are Blue... but Should Vqa Expect Them To?
- ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
- Efficient Probabilistic Logic Reasoning with Graph Neural Networks
- Clotho: An Audio Captioning Dataset
- Learn to Explain Efficiently via Neural Logic Inductive Learning
- SpatialSense: An Adversarially Crowdsourced Benchmark for Spatial Relation Recognition
- Probabilistic Logic Neural Networks for Reasoning
- OK-VQA: A Visual Question Answering Benchmark Requiring External\n Knowledge
- The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
- RAVEN: A Dataset for Relational and Analogical Visual rEasoNing
- A Corpus for Reasoning About Natural Language Grounded in Photographs
- Neural Ordinary Differential Equations
- Neural Probabilistic Logic Programming in DeepProbLog
- Trust-Aware Decision Making for Human-Robot Collaboration: Model Learning and Planning
- Vision-and-Language Navigation: Interpreting visually-grounded\n navigation instructions in real environments
- Attention Is All You Need
- Know-Evolve: Deep Temporal Reasoning for Dynamic Knowledge Graphs
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary\n Visual Reasoning
- Making the V in VQA Matter: Elevating the Role of Image Understanding in\n Visual Question Answering
- VQA: Visual Question Answering
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Long Short-Term Memory
- Finding Structure in Time
- SOAR: An architecture for general intelligence
- The Graph Neural Network Model
- An Integrative Theory of Prefrontal Cortex Function
- Gradient-based learning applied to document recognition
Cited by
Related