MiniCPM-V: A GPT-4V Level MLLM on Your Phone
2024/08/03 by Yuan Yao, Tianyou Yu, Yao, Yuan +41 · 207 citations
Computer Science · #Advanced Data Storage Technologies #Algorithms and Data Compression #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2408.01800
openalex publication_date 2024/08/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The recent surge of Multimodal Large Language Models (MLLMs) has fundamentally reshaped the landscape of AI research and industry, shedding light on a promising path toward the next AI milestone. However, significant challenges remain preventing MLLMs from being practical in real-world applications. The most notable challenge comes from the huge cost of running an MLLM with a massive number of parameters and extensive computation. As a result, most MLLMs need to be deployed on high-performing cloud servers, which greatly limits their application scopes such as mobile, offline, energy-sensitive, and privacy-protective scenarios. In this work, we present MiniCPM-V, a series of efficient MLLMs deployable on end-side devices. By integrating the latest MLLM techniques in architecture, pretraining and alignment, the latest MiniCPM-Llama3-V 2.5 has several notable features: (1) Strong performance, outperforming GPT-4V-1106, Gemini Pro and Claude 3 on OpenCompass, a comprehensive evaluation over 11 popular benchmarks, (2) strong OCR capability and 1.8M pixel high-resolution image perception at any aspect ratio, (3) trustworthy behavior with low hallucination rates, (4) multilingual support for 30+ languages, and (5) efficient deployment on mobile phones. More importantly, MiniCPM-V can be viewed as a representative example of a promising trend: The model sizes for achieving usable (e.g., GPT-4V) level performance are rapidly decreasing, along with the fast growth of end-side computation capacity. This jointly shows that GPT-4V level MLLMs deployed on end devices are becoming increasingly possible, unlocking a wider spectrum of real-world AI applications in the near future.
Cited by
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
- HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research
- Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
- Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models
- Benchmarking and Enhancing VLM for Compressed Image Understanding
- M3KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation
- Over++: Generative Video Compositing for Layer Interaction Effects
- QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning
- A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
- Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
- ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
- CitySeeker: How Do VLMS Explore Embodied Urban Navigation With Implicit Human Needs?
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- SuperCLIP: CLIP with Simple Classification Supervision
- AgentIAD: Agentic Industrial Anomaly Detection via Adaptive Memory Augmentation
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
- VisualActBench: Can VLMs See and Act like a Human?
- Towards Lossless Ultimate Vision Token Compression for VLMs
- Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels
- Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
- Jina-VLM: Small Multilingual Vision Language Model
- LoVoRA: Text-guided and Mask-free Video Object Removal and Addition with Learnable Object-aware Localization
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
- StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
- MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark
- ChartAnchor: Chart Grounding with Structural-Semantic Fidelity
- Charts Are Not Images: On the Challenges of Scientific Chart Editing
- DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering
- Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
- HKRAG: Holistic Knowledge Retrieval-Augmented Generation Over Visually-Rich Documents
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
- DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
- RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
- VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
- D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
- ChemLabs on ChemO: A Multi-Agent System for Multimodal Reasoning on IChO 2025
- MMD-Thinker: Adaptive Multi-Dimensional Thinking for Multimodal Misinformation Detection
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
- AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
- Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
- In-Orbit GRB Identification Using LLM-based model for the CXPD CubeSat
- Grounding Computer Use Agents on Human Demonstrations
- Federated Learning for Video Violence Detection: Complementary Roles of Lightweight CNNs and Vision-Language Models for Energy-Efficient Use
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures
- Unveiling Modality Bias: Automated Sample-Specific Analysis for Multimodal Misinformation Benchmarks
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- Towards Scalable Web Accessibility Audit with MLLMs as Copilots
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
- Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
- FOCUS: Efficient Keyframe Selection for Long Video Understanding
- GeoFM: Enhancing Geometric Reasoning of MLLMs via Synthetic Data Generation through Formal Language
- TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
- STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- Latent Chain-of-Thought for Visual Reasoning
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models
- MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- MedReason-R1: Learning to Reason for CT Diagnosis with Reinforcement Learning and Local Zoom
- MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
- SLICE: SLO-Driven Scheduling for LLM Inference on Edge Computing Devices
- Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents
- VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
- Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
- ReefNet: A Large scale, Taxonomically Enriched Dataset and Benchmark for Hard Coral Classification
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- Document Intelligence in the Era of Large Language Models: A Survey
- UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
- ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation
- A Survey on Agentic Multimodal Large Language Models
- Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
- Task-Specific Dual-Model Framework for Comprehensive Traffic Safety Video Description and Analysis
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- BitMar: Low-Bit Multimodal Fusion with Episodic Memory for Edge Devices
- TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
- Task-Aware Resolution Optimization for Visual Large Language Models
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
- Zero-shot image privacy classification with Vision-Language Models
- Towards Efficient Multimodal Unified Reasoning Model via Model Merging
- CapGeo: A Caption-Assisted Approach to Geometric Reasoning
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
- Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
- AgriGPT-VL: Agricultural Vision-Language Understanding Suite
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- OceanGym: A Benchmark Environment for Underwater Embodied Agents
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- NeMo: Needle in a Montage for Video-Language Understanding
- Falcon: A Cross-Modal Evaluation Dataset for Comprehensive Safety Perception
- Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning
- Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
- MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
- GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- Large AI Model-Enabled Generative Semantic Communications for Image Transmission
- MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
- QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding
- Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Live-E2T: Real-time Threat Monitoring in Video via Deduplicated Event Reasoning and Chain-of-Thought
- InstanceAssemble: Layout-Aware Image Generation via Instance Assembling Attention
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters
- TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding
- AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- ENJ: Optimizing Noise with Genetic Algorithms to Jailbreak LSMs
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- Towards Understanding Visual Grounding in Visual Language Models
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
- AdsQA: Towards Advertisement Video Understanding
- In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
- Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images
- Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
- LLM Enabled Multi-Agent System for 6G Networks: Framework and Method of Dual-Loop Edge-Terminal Collaboration
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
- Promptception: How Sensitive Are Large Multimodal Models to Prompts?
- E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
- VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
- KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
- PG-Agent: An Agent Powered by Page Graph
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- An Empirical Study on How Video-LLMs Answer Video Questions
- GM-Skip: Metric-Guided Transformer Block Skipping for Efficient Vision-Language Models
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- Mitigating Easy Option Bias in Multiple-Choice Question Answering
- edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer
- DianJin-OCR-R1: Enhancing OCR Capabilities via a Reasoning-and-Tool Interleaved Vision-Language Model
- E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- Region-Level Context-Aware Multimodal Understanding
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
- Causality Matters: How Temporal Information Emerges in Video Language Models
- UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
- The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
- OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
- SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
- STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
- Learning User Preferences for Image Generation Model
- Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
- Effective Training Data Synthesis for Improving MLLM Chart Understanding
- DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
- VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
- ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
- Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
- RealTalk-CN: A Realistic Chinese Speech-Text Dialogue Benchmark With Cross-Modal Interaction Analysis
- OmniPlay: Benchmarking Omni-Modal Models on Omni-Modal Game Playing
- ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
- Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacks
- GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
- SEA: Self-Evolution Agent with Step-wise Reward for Computer Use
- CogBench: A Large Language Model Benchmark for Multilingual Speech-Based Cognitive Impairment Assessment
- ADSeeker: A Knowledge-Infused Framework for Anomaly Detection and Reasoning
- Evaluating Variance in Visual Question Answering Benchmarks
- Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting
- Video-based Vehicle Surveillance in the Wild: License Plate, Make, and Model Recognition with Self Reflective Vision-Language Models
- Edge-Based Multimodal Sensor Data Fusion with Vision Language Models (VLMs) for Real-time Autonomous Vehicle Accident Avoidance
- Gems: Group Emotion Profiling Through Multimodal Situational Understanding
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
Related