MiniCPM-V: A GPT-4V Level MLLM on Your Phone
2024/08/03 by Yuan Yao, Tianyou Yu, Yao, Yuan +41 · 417 citations
Computer Science · #Advanced Data Storage Technologies #Algorithms and Data Compression #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2408.01800
openalex publication_date 2024/08/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The recent surge of Multimodal Large Language Models (MLLMs) has fundamentally reshaped the landscape of AI research and industry, shedding light on a promising path toward the next AI milestone. However, significant challenges remain preventing MLLMs from being practical in real-world applications. The most notable challenge comes from the huge cost of running an MLLM with a massive number of parameters and extensive computation. As a result, most MLLMs need to be deployed on high-performing cloud servers, which greatly limits their application scopes such as mobile, offline, energy-sensitive, and privacy-protective scenarios. In this work, we present MiniCPM-V, a series of efficient MLLMs deployable on end-side devices. By integrating the latest MLLM techniques in architecture, pretraining and alignment, the latest MiniCPM-Llama3-V 2.5 has several notable features: (1) Strong performance, outperforming GPT-4V-1106, Gemini Pro and Claude 3 on OpenCompass, a comprehensive evaluation over 11 popular benchmarks, (2) strong OCR capability and 1.8M pixel high-resolution image perception at any aspect ratio, (3) trustworthy behavior with low hallucination rates, (4) multilingual support for 30+ languages, and (5) efficient deployment on mobile phones. More importantly, MiniCPM-V can be viewed as a representative example of a promising trend: The model sizes for achieving usable (e.g., GPT-4V) level performance are rapidly decreasing, along with the fast growth of end-side computation capacity. This jointly shows that GPT-4V level MLLMs deployed on end devices are becoming increasingly possible, unlocking a wider spectrum of real-world AI applications in the near future.
Cited by
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
- HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research
- Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
- Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models
- Benchmarking and Enhancing VLM for Compressed Image Understanding
- M3KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation
- Over++: Generative Video Compositing for Layer Interaction Effects
- QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning
- A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
- Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
- ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
- CitySeeker: How Do VLMS Explore Embodied Urban Navigation With Implicit Human Needs?
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- SuperCLIP: CLIP with Simple Classification Supervision
- AgentIAD: Agentic Industrial Anomaly Detection via Adaptive Memory Augmentation
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
- VisualActBench: Can VLMs See and Act like a Human?
- Towards Lossless Ultimate Vision Token Compression for VLMs
- Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels
- Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
- Jina-VLM: Small Multilingual Vision Language Model
- LoVoRA: Text-guided and Mask-free Video Object Removal and Addition with Learnable Object-aware Localization
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
- StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
- MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark
- ChartAnchor: Chart Grounding with Structural-Semantic Fidelity
- Charts Are Not Images: On the Challenges of Scientific Chart Editing
- DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering
- Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
- HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
- DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
- RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
- VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
- D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
- ChemLabs on ChemO: A Multi-Agent System for Multimodal Reasoning on IChO 2025
- MMD-Thinker: Adaptive Multi-Dimensional Thinking for Multimodal Misinformation Detection
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
- AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
- Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
- In-Orbit GRB Identification Using LLM-based model for the CXPD CubeSat
- Grounding Computer Use Agents on Human Demonstrations
- Federated Learning for Video Violence Detection: Complementary Roles of Lightweight CNNs and Vision-Language Models for Energy-Efficient Use
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures
- Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- Towards Scalable Web Accessibility Audit with MLLMs as Copilots
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
- Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
- FOCUS: Efficient Keyframe Selection for Long Video Understanding
- GeoFM: Enhancing Geometric Reasoning of MLLMs via Synthetic Data Generation through Formal Language
- TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
- STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- Latent Chain-of-Thought for Visual Reasoning
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models
- MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- MedReason-R1: Learning to Reason for CT Diagnosis with Reinforcement Learning and Local Zoom
- MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
- SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
- Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents
- VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
- Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
- ReefNet: A Large-Scale Dataset and Benchmark for Fine-Grained Coral Reef Recognition
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- Document Intelligence in the Era of Large Language Models: A Survey
- UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
- ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
- A Survey on Agentic Multimodal Large Language Models
- Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
- Task-Specific Dual-Model Framework for Comprehensive Traffic Safety Video Description and Analysis
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- BitMar: Low-Bit Multimodal Fusion with Episodic Memory for Edge Devices
- TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
- Task-Aware Resolution Optimization for Visual Large Language Models
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
- Zero-shot image privacy classification with Vision-Language Models
- Towards Efficient Multimodal Unified Reasoning Model via Model Merging
- CapGeo: A Caption-Assisted Approach to Geometric Reasoning
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
- Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
- AgriGPT-VL: Agricultural Vision-Language Understanding Suite
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- OceanGym: A Benchmark Environment for Underwater Embodied Agents
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- NeMo: Needle in a Montage for Video-Language Understanding
- Falcon: A Cross-Modal Evaluation Dataset for Comprehensive Safety Perception
- Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning
- Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
- MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
- GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- Large AI Model-Enabled Generative Semantic Communications for Image Transmission
- MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
- QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding
- Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Live-E2T: Real-time Threat Monitoring in Video via Deduplicated Event Reasoning and Chain-of-Thought
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
- InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters
- TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding
- AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- ERIS: Evolutionary Real-world Interference Scheme for Jailbreaking Audio Large Models
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- Towards Understanding Visual Grounding in Visual Language Models
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
- AdsQA: Towards Advertisement Video Understanding
- In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
- Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images
- Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
- LLM Enabled Multi-Agent System for 6G Networks: Framework and Method of Dual-Loop Edge-Terminal Collaboration
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
- Promptception: How Sensitive Are Large Multimodal Models to Prompts?
- E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
- VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
- KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
- PG-Agent: An Agent Powered by Page Graph
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- An Empirical Study on How Video-LLMs Answer Video Questions
- GM-Skip: Metric-Guided Transformer Block Skipping for Efficient Vision-Language Models
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- Mitigating Easy Option Bias in Multiple-Choice Question Answering
- edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer
- DianJin-OCR-R1: Enhancing OCR Capabilities via a Reasoning-and-Tool Interleaved Vision-Language Model
- E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- Region-Level Context-Aware Multimodal Understanding
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
- Causality Matters: How Temporal Information Emerges in Video Language Models
- UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
- The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
- OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
- SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
- STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
- Learning User Preferences for Image Generation Model
- Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
- Effective Training Data Synthesis for Improving MLLM Chart Understanding
- DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
- VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
- ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
- Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
- RealTalk-CN: A Realistic Chinese Speech-Text Dialogue Benchmark With Cross-Modal Interaction Analysis
- OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
- ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
- Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacks
- GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
- SEA: Self-Evolution Agent with Step-wise Reward for Computer Use
- CogBench: A Large Language Model Benchmark for Multilingual Speech-Based Cognitive Impairment Assessment
- ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and Reasoning
- Evaluating Variance in Visual Question Answering Benchmarks
- Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting
- Video-based Vehicle Surveillance in the Wild: License Plate, Make, and Model Recognition with Self Reflective Vision-Language Models
- Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance
- VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
- Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
- Gems: Group Emotion Profiling Through Multimodal Situational Understanding
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
- In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
- LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
- LMM-Det: Make Large Multimodal Models Excel in Object Detection
- U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
- MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
- Docopilot: Improving Multimodal Models for Document-Level Understanding
- C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
- AD2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions
- AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
- AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
- Describe Anything Model for Visual Question Answering on Text-rich Images
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding
- MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
- DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations
- Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
- MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
- UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
- FaceLLM: A Multimodal Large Language Model for Face Understanding
- Multilingual Multimodal Software Developer for Code Generation
- Visual Semantic Description Generation with MLLMs for Image-Text Matching
- MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation
- SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs
- JoyTTS: LLM-based Spoken Chatbot With Voice Cloning
- CREST: Curvature-Regulated Event-Centric Sampling for Efficient Long-Video Understanding
- Kwai Keye-VL Technical Report
- Multimodal Chip Physical Design Engineer Assistant
- CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs
- Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
- From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
- WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
- SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding
- HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
- Team PA-VCG's Solution for Competition on Understanding Chinese College Entrance Exam Papers in ICDAR'25
- FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
- Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
- Synthetic Visual Genome
- Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
- Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition
- MiniCPM4: Ultra-Efficient LLMs on End Devices
- AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown
- WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code
- MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
- Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
- SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
- A Neurosymbolic Agent System for Compositional Visual Reasoning
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generation
- MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
- Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
- AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
- FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design
- AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
- EraserDiT: Fast Video Inpainting with Diffusion Transformer Model
- LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
- SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models
- Exploring the Secondary Risks of Large Language Models
- Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
- DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning
- CogStream: Context-guided Streaming Video Question Answering
- DreamCS: Geometry-Aware Text-to-3D Generation with Unpaired 3D Reward Supervision
- Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
- CoMemo: LVLMs Need Image Context with Image Memory
- SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs
- AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
- Multimodal Tabular Reasoning with Privileged Structured Information
- SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
- Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
- DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
- LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
- Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
- Native-Resolution Image Synthesis
- CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
- Seeing the Arrow of Time in Large Multimodal Models
- Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
- SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
- AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning
- Flow2Code: Evaluating Large Language Models for Flowchart-based Code Generation Capability
- MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
- Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
- FLEX: A Largescale Multimodal, Multiview Dataset for Learning Structured Representations for Fitness Action Quality Assessment
- ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
- Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing
- VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
- Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
- Affordance Benchmark for MLLMs
- GuessBench: Sensemaking Multimodal Creativity in the Wild
- SynSHRP2: A Synthetic Multimodal Benchmark for Driving Safety-critical Events Derived from Real-world Driving Data
- Ivy-Fake: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection
- Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
- SA-Person: Text-Based Person Retrieval with Scene-aware Re-ranking
- DisTime: Distribution-based Time Representation for Video Large Language Models
- Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
- Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
- Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
- Mobi-π: Mobilizing Your Robot Learning Policy
- Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information
- VModA: An Effective Framework for Adaptive NSFW Image Moderation
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
- VidText: Towards Comprehensive Evaluation for Video Text Understanding
- OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
- AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
- TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
- TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
- MLLM-Guided VLM Fine-Tuning with Joint Inference for Zero-Shot Composed Image Retrieval
- USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models
- Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
- Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
- Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
- Generative RLHF-V: Learning Principles from Multi-modal Human Preference
- HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving
- FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain
- RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
- GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
- Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
- R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
- Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
- ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation
- From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
- STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
- MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
- Clapper: Compact Learning and Video Representation in VLMs
- Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents
- UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
- Beyond Text: Unveiling Privacy Vulnerabilities in Multi-modal Retrieval-Augmented Generation
- Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
- Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
- Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
- From Shots to Stories: LLM-Assisted Video Editing with Unified Language Representations
- LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
- SafeVid: Toward Safety Aligned Video Large Multimodal Models
- Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
- MedSG-Bench: A Benchmark for Medical Image Sequences Grounding
- WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild?
- PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language
- MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
- Emotion Knowledge Enhancement for Vision Large Language Models: A Self-Verification Approach for High-Quality Emotion Instruction Data Generation
- Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
- Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
- Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning
- Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
- Neural Catalog: Scaling Species Recognition with Catalog of Life-Augmented Generation
- FG-CLIP: Fine-Grained Visual and Textual Alignment
- LecEval: An Automated Metric for Multimodal Knowledge Acquisition in Multimedia Learning
- Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
- Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
- OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Models
- Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering
- SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
- LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs
- Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- Zero-shot adaptable task planning for autonomous construction robots: a comparative study of lightweight single and multi-AI agent systems
- Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
- OmniGUI: Benchmarking GUI Agents in Omni-Modal Smartphone Environments
- OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models
- VEU-Bench: Towards Comprehensive Understanding of Video Editing
- DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models
- MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
- MOOSComp: Improving Lightweight Long-Context Compressor via Mitigating Over-Smoothing and Incorporating Outlier Scores
- Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos
- Towards Visual Text Grounding of Multimodal Large Language Model
- LLM-Alignment Live-Streaming Recommendation
- VideoVista-CulturalLingo: 360^∘ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
- Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
- Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
- FaceInsight: A Multimodal Large Language Model for Face Perception
- LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
- CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
- An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding
- Manipulating Multimodal Agents via Cross-Modal Prompt Injection
- VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
- LAD-Reasoner: Tiny Multimodal Models are Good Reasoners for Logical Anomaly Detection
- FashionDPO:Fine-tune Fashion Outfit Generation Model using Direct Preference Optimization
- Evaluating Menu OCR and Translation: A Benchmark for Aligning Human and Automated Evaluations in Large Vision-Language Models
- Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
- AnomalyR1: A GRPO-based End-to-end MLLM for Industrial Anomaly Detection
- OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
- MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- PRM-BAS: Enhancing Multimodal Reasoning through PRM-guided Beam Annealing Search
- Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
- Resampling Benchmark for Efficient Comprehensive Evaluation of Large Vision-Language Models
- VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
- VideoAds for Fast-Paced Video Understanding
- LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
- MM-IFEngine: Towards Multimodal Instruction Following
- VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
- SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
- OmniCaptioner: One Captioner to Rule Them All
- The Human Robot Social Interaction (HSRI) Dataset: Benchmarking Foundational Models' Social Reasoning
- SmolVLM: Redefining small and efficient multimodal models
Related