MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
2025/08/03 by Yi Liu, Liu, Yi, Xu Xiao +20
Computer Science · #Multimodal Machine Learning Applications #Advanced Neural Network Applications #ICT in Developing Communities
paper · pdf · doi:10.48550/arxiv.2508.01540
Abstract
Vision-Language Models (VLMs) have achieved remarkable breakthroughs in recent years, enabling a diverse array of applications in everyday life. However, the substantial computational and storage demands of VLMs pose significant challenges for their efficient deployment on mobile devices, which represent the most ubiquitous and accessible computing platforms today. In this work, we introduce MagicVL-2B, a novel VLM meticulously optimized for flagship smartphones. MagicVL-2B leverages a lightweight visual encoder with fewer than 100M parameters and features a redesigned dynamic resolution scheme that adaptively generates image tokens without excessive modification of image dimensions. To further enhance the performance of this compact encoder within VLMs, we propose a multimodal curriculum learning strategy that incrementally increases task difficulty and data information density throughout training. This approach substantially improves the model's performance across a variety of sub-tasks. Extensive evaluations on standard VLM benchmarks demonstrate that MagicVL-2B matches the accuracy of current state-of-the-art models while reducing on-device power consumption by 41.1%. These results establish MagicVL-2B as a practical and robust solution for real-world mobile vision-language applications, enabling advanced multimodal intelligence to run directly on smartphones.
Citations
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- SmolVLM: Redefining small and efficient multimodal models
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- NVILA: Efficient Frontier Visual Language Models
- BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
- Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
- Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
- MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
- LLaVA-OneVision: Easy Visual Task Transfer
- Mobile Edge Intelligence for Large Language Models: A Contemporary Survey
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- PowerInfer-2: Fast Large Language Model Inference on a Smartphone
- Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- OpenELM: An Efficient Language Model Family with Open Training and Inference Framework
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- Transformer-Lite: High-efficiency Deployment of Large Language Models on Mobile Phone GPUs
- MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
- Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- VILA: On Pre-training for Visual Language Models
- Honeybee: Locality-enhanced Projector for Multimodal LLM
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
- Sigmoid Loss for Language Image Pre-Training
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- LLaMA: Open and Efficient Foundation Language Models
- Learning Transferable Visual Models From Natural Language Supervision
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Language Models are Few-Shot Learners
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Related