2024/02/06 by Xiangxiang Chu, Limeng Qiao, Chu, Xiangxiang +19 · 21 citations
Computer Science · Engineering · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Robotics and Automated Systems
paper · pdf · doi:10.48550/arxiv.2402.03766
We introduce MobileVLM V2, a family of significantly improved vision language models upon MobileVLM, which proves that a delicate orchestration of novel architectural design, an improved training scheme tailored for mobile VLMs, and rich high-quality dataset curation can substantially benefit VLMs' performance. Specifically, MobileVLM V2 1.7B achieves better or on-par performance on standard VLM benchmarks compared with much larger VLMs at the 3B scale. Notably, our 3B model outperforms a large variety of VLMs at the 7B+ scale. Our models will be released at https://github.com/Meituan-AutoML/MobileVLM .