A Survey on Agentic Multimodal Large Language Models
2025/10/13 by Yao, Huanjin, Zhang, Ruifei, Huang, Jiaxing +8 · 4 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.10991
Abstract
With the recent emergence of revolutionary autonomous agentic systems, research community is witnessing a significant shift from traditional static, passive, and domain-specific AI agents toward more dynamic, proactive, and generalizable agentic AI. Motivated by the growing interest in agentic AI and its potential trajectory toward AGI, we present a comprehensive survey on Agentic Multimodal Large Language Models (Agentic MLLMs). In this survey, we explore the emerging paradigm of agentic MLLMs, delineating their conceptual foundations and distinguishing characteristics from conventional MLLM-based agents. We establish a conceptual framework that organizes agentic MLLMs along three fundamental dimensions: (i) Agentic internal intelligence functions as the system's commander, enabling accurate long-horizon planning through reasoning, reflection, and memory; (ii) Agentic external tool invocation, whereby models proactively use various external tools to extend their problem-solving capabilities beyond their intrinsic knowledge; and (iii) Agentic environment interaction further situates models within virtual or physical environments, allowing them to take actions, adapt strategies, and sustain goal-directed behavior in dynamic real-world scenarios. To further accelerate research in this area for the community, we compile open-source training frameworks, training and evaluation datasets for developing agentic MLLMs. Finally, we review the downstream applications of agentic MLLMs and outline future research directions for this rapidly evolving field. To continuously track developments in this rapidly evolving field, we will also actively update a public repository at https://github.com/HJYao00/Awesome-Agentic-MLLMs.
Citations
- FineVision: Open Data Is All You Need
- Towards Efficient Multimodal Unified Reasoning Model via Model Merging
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- MAPO: Mixed Advantage Policy Optimization
- Qwen3-Omni Technical Report
- InfraMind: A Novel Exploration-based GUI Agentic Framework for Mission-critical Industrial Management
- Scaling Agents via Continual Pre-training
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
- Igniting VLMs toward the Embodied Space
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
- Nav-R1: Reasoning and Navigation in Embodied Scenes
- RecoWorld: Building Simulated Environments for Agentic Recommender Systems
- OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning
- MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
- A Survey of Reinforcement Learning for Large Reasoning Models
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- Long-Horizon Visual Imitation Learning via Plan and Code Reflection
- The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
- AutoDrive-R2: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Kwai Keye-VL 1.5 Technical Report
- Reinforced Visual Perception with Tools
- EO-1: An Open Unified Embodied Foundation Model for General Robot Control
- Embodied AI: Emerging Risks and Opportunities for Policy Action
- rStar2-Agent: Agentic Reasoning Technical Report
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- AT-CXR: Uncertainty-Aware Agentic Triage for Chest X-rays
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Towards Better Correctness and Efficiency in Code Generation
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
- Mobile-Agent-v3: Fundamental Agents for GUI Automation
- MedResearcher-R1: Expert-Level Medical Deep Researcher via A Knowledge-Informed Trajectory Synthesis Framework
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- OS-R1: Agentic Operating System Kernel Tuning with Reinforcement Learning
- Thyme: Think Beyond Images
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- MolmoAct: Action Reasoning Models that can Reason in Space
- gpt-oss-120b & gpt-oss-20b Model Card
- M2IO-R1: An Efficient RL-Enhanced Reasoning Framework for Multimodal Retrieval Augmented Multimodal Generation
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- Posterior-GRPO: Rewarding Reasoning Processes in Code Generation
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
- VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
- MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations
- Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding
- InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
- GR-3 Technical Report
- WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
- Scaling RL to Long Videos
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
- WebSailor: Navigating Super-human Reasoning for Web Agent
- Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
- SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
- Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
- MMSearch-R1: Incentivizing LMMs to Search
- Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning
- Inference-Time Reward Hacking in Large Language Models
- MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
- From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
- Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
- Deep Research Agents: A Systematic Examination And Roadmap
- VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
- GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
- AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need
- MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
- ComfyUI-R1: Exploring Reasoning Models for Workflow Generation
- OctoNav: Towards Generalist Embodied Navigation
- DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts
- CoRT: Code-integrated Reasoning within Thinking
- FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
- Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
- VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
- Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
- TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems
- MiMo-VL Technical Report
- Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
- Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning
- MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
- SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
- GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
- ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
- WebDancer: Towards Autonomous Information Seeking Agency
- Reinforced Reasoning for Embodied Planning
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
- Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models
- Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
- rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset
- Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
- MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability
- WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
- R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search
- Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
- GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
- Training-Free Reasoning and Reflection in MLLMs
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
- R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
- AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
- VLM-R3: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
- AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
- Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- Visual Agentic Reinforcement Fine-Tuning
- Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
- AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning
- Visual Planning: Let's Think Only with Images
- Qwen3 Technical Report
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
- Seed1.5-VL Technical Report
- UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
- EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
- RM-R1: Reward Modeling as Reasoning
- DriveAgent: Multi-Agent Structured Reasoning with LLM and Multimodal Sensor Fusion for Autonomous Driving
- Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
- π0.5: a Vision-Language-Action Model with Open-World Generalization
- InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
- TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
- NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- A Survey of Frontiers in LLM Reasoning: Inference Scaling, Learning to Reason, and Agentic Systems
- VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
- SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- Kimi-VL Technical Report
- SmolVLM: Redefining small and efficient multimodal models
- DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments
- ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
- ToRL: Scaling Tool-Integrated RL
- Video-R1: Reinforcing Video Reasoning in MLLMs
- Qwen2.5-Omni Technical Report
- Gemma 3 Technical Report
- Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding
- Enhancing Vision Foundation Models via Multimodal Continual Pre-Training
- Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
- Towards Agentic Recommender Systems in the Era of Multimodal Large Language Models
- UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
- R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
- Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
- VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
- DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
- Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
- R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents
- AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning
- Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
- LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
- GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images
- R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
- Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model
- Advancing vision-language models in front-end development via data synthesis
- CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering
- ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
- Scalable Best-of-N Selection for Large Language Models via Self-Certainty
- Qwen2.5-VL Technical Report
- Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical Knowledge
- A-MEM: Agentic Memory for LLM Agents
- VLP: Vision-Language Preference Learning for Embodied Manipulation
- MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
- ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models
- WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point
- AStar: Boosting Multimodal Reasoning with Automated Structured Thinking
- ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
- Process-Supervised Reinforcement Learning for Code Generation
- M+: Extending MemoryLLM with Scalable Long-Term Memory
- Baichuan-Omni-1.5 Technical Report
- Humanity's Last Exam
- Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
- O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- UI-TARS: Pioneering Automated GUI Interaction with Native Agents
- Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
- WebWalker: Benchmarking LLMs in Web Traversal
- Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
- Search-o1: Agentic Search-Enhanced Large Reasoning Models
- InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
- rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
- Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning
- Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
- Phi-4 Technical Report
- MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
- Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
- ShowUI: One Vision-Language-Action Model for GUI Visual Agent
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
- AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations
- Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
- Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
- π0: A Vision-Language-Action Flow Model for General Robot Control
- Vision-Language Models Can Self-Improve Reasoning via Reflection
- Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
- GPT-4o System Card
- SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models
- Improve Vision Language Model Chain-of-thought Reasoning
- MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
- Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models
- MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
- MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
- RepoGenReflex: Enhancing Repository-Level Code Completion with Verbal Reinforcement and Retrieval-Augmented Generation
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
- PhishAgent: A Robust Multimodal Agent for Phishing Webpage Detection
- LongVILA: Scaling Long-Context Visual Language Models for Long Videos
- Harnessing Multimodal Large Language Models for Multimodal Sequential Recommendation
- LLaVA-OneVision: Easy Visual Task Transfer
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- The Llama 3 Herd of Models
- MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
- MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
- AutoFlow: Automated Workflow Generation for Large Language Model Agents
- HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
- CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
- Long Context Transfer from Language to Vision
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
- OpenVLA: An Open-Source Vision-Language-Action Model
- LVBench: An Extreme Long Video Understanding Benchmark
- GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices
- MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models
- On the Effects of Data Scale on UI Control Agents
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
- M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
- Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
- Automated Multi-level Preference for MLLMs
- What matters when building vision-language models?
- MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
- PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
- TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding
- Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
- Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
- LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
- mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
- VideoAgent: Long-form Video Understanding with Large Language Model as Agent
- Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
- Investigating Continual Pretraining in Large Language Models: Insights and Implications
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
- Large Multimodal Agents: A Survey
- Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
- AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
- Rec-GPT4V: Multimodal Recommendation with Large Vision-Language Models
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
- StepCoder: Improve Code Generation with Reinforcement Learning from Compiler Feedback
- Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios
- MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
- InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
- LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
- ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation
- V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs
- MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
- Compositional Chain-of-Thought Prompting for Large Multimodal Models
- GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
- Evidential Active Recognition: Intelligent and Prudent Open-World Embodied Perception
- Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
- HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs
- LLM4Drive: A Survey of Large Language Models for Autonomous Driving
- MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
- Improved Baselines with Visual Instruction Tuning
- MetaTool Benchmark for Large Language Models: Deciding Whether to Use Tools and Which to Use
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
- NExT-GPT: Any-to-Any Multimodal LLM
- EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
- Continual Pre-Training of Large Language Models: How to (re)warm your model?
- MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
- MMBench: Is Your Multi-modal Model an All-around Player?
- LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
- ALP: Action-Aware Embodied Learning for Perception
- Valley: Video Assistant with Large Language model Enhanced abilitY
- Mind2Web: Towards a Generalist Agent for the Web
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- PaCE: Unified Multi-modal Dialogue Pre-training with Progressive and Compositional Experts
- HuatuoGPT, towards Taming Language Model to Be a Doctor
- MemoryBank: Enhancing Large Language Models with Long-Term Memory
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Visual Instruction Tuning
- Vision-Language Models for Vision Tasks: A Survey
- Vision-Language Models for Vision Tasks: A Survey
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- Multimodal Chain-of-Thought Reasoning in Language Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- REVEAL: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge Memory
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and Text
- CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
- Flamingo: a Visual Language Model for Few-Shot Learning
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- WebGPT: Browser-assisted question-answering with human feedback
- Safe Reinforcement Learning using Formal Verification for Tissue Retraction in Autonomous Robotic-Assisted Surgery
- Evaluating Large Language Models Trained on Code
- Learning to Explore using Active Neural SLAM
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Neural Modular Control for Embodied Question Answering
- Learning to Look Around: Intelligently Exploring Unseen Environments for\n Unknown Tasks
- Proximal Policy Optimization Algorithms
- R1-Code-Interpreter: LLMs Reason with Code via Supervised and Multi-stage Reinforcement Learning
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
- MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
- MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents
- Acting Less is Reasoning More! Teaching Model to Act Efficiently
- Gemini Robotics: Bringing AI into the Physical World
- HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- OpenAI o1 System Card
Cited by
Related