DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
2024/12/13 by Zhiyu Wu, Wu, Zhiyu, Xiaokang Chen +51 · 3 voices · 140 citations
Computer Science · #cs.CV #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2412.10302
Abstract
We present DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL, through two key major upgrades. For the vision component, we incorporate a dynamic tiling vision encoding strategy designed for processing high-resolution images with different aspect ratios. For the language component, we leverage DeepSeekMoE models with the Multi-head Latent Attention mechanism, which compresses Key-Value cache into latent vectors, to enable efficient inference and high throughput. Trained on an improved vision-language dataset, DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question answering, optical character recognition, document/table/chart understanding, and visual grounding. Our model series is composed of three variants: DeepSeek-VL2-Tiny, DeepSeek-VL2-Small and DeepSeek-VL2, with 1.0B, 2.8B and 4.5B activated parameters respectively. DeepSeek-VL2 achieves competitive or state-of-the-art performance with similar or fewer activated parameters compared to existing open-source dense and MoE-based models. Codes and pre-trained models are publicly accessible at https://github.com/deepseek-ai/DeepSeek-VL2.
Cited by
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratory
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- Group Preference Collapse in Personalized Multimodal Large Language Models
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Semantic Mismatch and Perceptual Degradation: A New Perspective on Image Editing Immunity
- MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data
- DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry
- CAPTURE: A Benchmark and Evaluation for LVLMs in CAPTCHA Resolving
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
- What really matters for person re-identification? A Mixture-of-Experts Framework for Semantic Attribute Importance
- Beyond Real Weights: Hypercomplex Representations for Stable Quantization
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- VisKnow: Constructing Visual Knowledge Base for Object Understanding
- NeuroABench: A Multimodal Evaluation Benchmark for Neurosurgical Anatomy Identification
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
- VRSA: Jailbreaking Multimodal Large Language Models through Visual Reasoning Sequential Attack
- BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models
- UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
- Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
- Hierarchical Process Reward Models are Symbolic Vision Learners
- Lost in Modality: Evaluating the Effectiveness of Text-Based Membership Inference Attacks on Large Multimodal Models
- VACoT: Rethinking Visual Data Augmentation with VLMs
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- Seeing before Observable: Potential Risk Reasoning in Autonomous Driving via Vision Language Models
- RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
- Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
- Benchmarking Corruption Robustness of LVLMs: A Discriminative Benchmark and Robustness Alignment Metric
- Disc3D: Automatic Curation of High-Quality 3D Dialog Data via Discriminative Object Referring
- ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
- FastMMoE: Accelerating Multimodal Large Language Models through Dynamic Expert Activation and Routing-Aware Token Pruning
- ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
- MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models
- O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
- P1: Mastering Physics Olympiads with Reinforcement Learning
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
- TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
- SRNN: Spatiotemporal Relational Neural Network for Intuitive Physics Understanding
- Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption Supervision
- Route Experts by Sequence, not by Token
- Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
- V-Thinker: Interactive Thinking with Images
- What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
- Seeing, Signing, and Saying: A Vision-Language Model-Assisted Pipeline for Sign Language Data Acquisition and Curation from Social Media
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
- DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry
- Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
- Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- Automated urban waterlogging assessment and early warning through a mixture of foundation models
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- DeepSeek-OCR: Contexts Optical Compression
- Token-Level Inference-Time Alignment for Vision-Language Models
- Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
- Region in Context: Text-condition Image editing with Human-like semantic reasoning
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- Detect Anything via Next Point Prediction
- SpineBench: Benchmarking Multimodal LLMs for Spinal Pathology Analysis
- Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model
- A Survey on Agentic Multimodal Large Language Models
- MC#: Mixture Compressor for Mixture-of-Experts Large Models
- Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
- Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving
- PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- FinMR: A Knowledge-Intensive Multimodal Benchmark for Advanced Financial Reasoning
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
- LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
- SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
- RServe: Overlapping Encoding and Prefill for Efficient LMM Inference
- VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
- DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
- MMPB: It's Time for Multi-Modal Personalization
- UrbanFeel: A Comprehensive Benchmark for Temporal and Perceptual Understanding of City Scenes through Human Perspective
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning
- PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology
- F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model
- VAInpaint: Zero-Shot Video-Audio inpainting framework with LLMs-driven Module
- Can GRPO Boost Complex Multimodal Table Understanding?
- MoE-Inference-Bench: Performance Evaluation of Mixture of Expert Large Language and Vision Models
- A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- AsyMoE: Leveraging Modal Asymmetry for Enhanced Expert Specialization in Large Vision-Language Models
- Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents
- AesBiasBench: Evaluating Bias and Alignment in Multimodal Language Models for Personalized Image Aesthetic Assessment
- PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments
- Measuring Epistemic Humility in Multimodal Large Language Models
- Multimodal LLMs See Sentiment
- Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- MoPEQ: Mixture of Mixed Precision Quantized Experts
- LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
- The Demon is in Ambiguity: Revisiting Situation Recognition with Single Positive Multi-Label Learning
- Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- Estimating 2D Keypoints of Surgical Tools Using Vision-Language Models with Low-Rank Adaptation
- SUMMA: A Multimodal Large Language Model for Advertisement Summarization
- OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
- KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- Scene-Aware Vectorized Memory Multi-Agent Framework with Cross-Modal Differentiated Quantization VLMs for Visually Impaired Assistance
- MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs
- Mitigating Easy Option Bias in Multiple-Choice Question Answering
- RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- The Perils of Chart Deception: How Misleading Visualizations Affect Vision-Language Models
- SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
- SOI is the Root of All Evil: Quantifying and Breaking Similar Object Interference in Single Object Tracking
- Multimodal Recommendation via Self-Corrective Preference Alignmen
- Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
- mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering
- Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
- PET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography
- Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
- SpectrumWorld: Artificial Intelligence Foundation for Spectroscopy
- Oedipus and the Sphinx: Benchmarking and Improving Visual Language Models for Complex Graphic Reasoning
- CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning
- AutoBridge: Automating Smart Device Integration with Centralized Platform
- MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models
- A Large Language Model Powered Integrated Circuit Footprint Geometry Understanding
- MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
- TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
- Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
- RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
Discussions
Related