Jina-VLM: Small Multilingual Vision Language Model
2025/12/03 by Koukounas, Andreas, Mastrapas, Georgios, Hönicke, Florian +4
#68T50 #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #I.2.7
paper · doi:10.48550/arxiv.2512.04032
Abstract
We present Jina-VLM, a 2.4B parameter vision-language model that achieves state-of-the-art multilingual visual question answering among open 2B-scale VLMs. The model couples a SigLIP2 vision encoder with a Qwen3 language backbone through an attention-pooling connector that enables token-efficient processing of arbitrary-resolution images. The model achieves leading results on standard VQA benchmarks and multilingual evaluations while preserving competitive text-only performance. Model weights and code are publicly released at https://huggingface.co/jinaai/jina-vlm .
Citations
- FineVision: Open Data Is All You Need
- Multilingual Vision-Language Models, A Survey
- HERO: Rethinking Visual Token Early Dropping in High-Resolution Large Vision-Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Ovis2.5 Technical Report
- Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
- Qwen3 Technical Report
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- SmolVLM: Redefining small and efficient multimodal models
- VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- NVILA: Efficient Frontier Visual Language Models
- VisionZip: Longer is Better but Not Necessary in Vision Language Models
- PaliGemma 2: A Family of Versatile VLMs for Transfer
- PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
- Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
- R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- LLaVA-OneVision: Easy Visual Task Transfer
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
- PaliGemma: A versatile 3B VLM for transfer
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
- We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
- Parrot: Multilingual Visual Instruction Tuning
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering
- Imp: Highly Capable Large Multimodal Models for Mobile Devices
- What matters when building vision-language models?
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
- MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
- Rotary Position Embedding for Vision Transformer
- LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
- An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
- Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
- Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning
- MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
- Improved Baselines with Visual Instruction Tuning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- MMBench: Is Your Multi-modal Model an All-around Player?
- Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Evaluating Object Hallucination in Large Vision-Language Models
- Visual Instruction Tuning
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots
- PaLI: A Jointly-Scaled Multilingual Language-Image Model
- Flamingo: a Visual Language Model for Few-Shot Learning
- ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
- Training Verifiers to Solve Math Word Problems
- TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance
- InfographicVQA
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Measuring Massive Multitask Language Understanding
- DocVQA: A Dataset for VQA on Document Images
- PathVQA: 30000+ Questions for Medical Visual Question Answering
- GLU Variants Improve Transformer
- HellaSwag: Can a Machine Really Finish Your Sentence?
- Towards VQA Models That Can Read
- TallyQA: Answering Complex Counting Questions
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- A Diagram Is Worth A Dozen Images
Related