VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
2024/06/12 by Jiannan Wu, Wu, Jiannan, Muyan Zhong +23 · 68 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2406.08394
openalex publication_date 2024/06/12 · openalex created_date 2024/06/14 · openalex updated_date 2026/07/28
Abstract
We present VisionLLM v2, an end-to-end generalist multimodal large model (MLLM) that unifies visual perception, understanding, and generation within a single framework. Unlike traditional MLLMs limited to text output, VisionLLM v2 significantly broadens its application scope. It excels not only in conventional visual question answering (VQA) but also in open-ended, cross-domain vision tasks such as object localization, pose estimation, and image generation and editing. To this end, we propose a new information transmission mechanism termed "super link", as a medium to connect MLLM with task-specific decoders. It not only allows flexible transmission of task information and gradient feedback between the MLLM and multiple downstream decoders but also effectively resolves training conflicts in multi-tasking scenarios. In addition, to support the diverse range of tasks, we carefully collected and combed training data from hundreds of public vision and vision-language tasks. In this way, our model can be joint-trained end-to-end on hundreds of vision language tasks and generalize to these tasks using a set of shared parameters through different user prompts, achieving performance comparable to task-specific models. We believe VisionLLM v2 will offer a new perspective on the generalization of MLLMs.
Cited by
- Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
- Multi-Grained Text-Guided Image Fusion for Multi-Exposure and Multi-Focus Scenarios
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Moment and Highlight Detection via MLLM Frame Segmentation
- SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
- Grounding Everything in Tokens for Multimodal Large Language Models
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
- Beyond Real versus Fake Towards Intent-Aware Video Analysis
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- Co-Training Vision Language Models for Remote Sensing Multi-task Learning
- V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs
- DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
- Draft and Refine with Visual Experts
- Sim4Seg: Boosting Multimodal Multi-disease Medical Diagnosis Segmentation with Region-Aware Vision-Language Similarity Masks
- PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
- LLM-as-a-Judge for Software Engineering: Literature Review, Vision, and the Road Ahead
- SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- Token-Level Inference-Time Alignment for Vision-Language Models
- Detect Anything via Next Point Prediction
- MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding
- Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Falcon: A Cross-Modal Evaluation Dataset for Comprehensive Safety Perception
- Guiding Audio Editing with Audio Language Model
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- VideoAVE: A Multi-Attribute Video-to-Text Attribute Value Extraction Dataset and Benchmark Models
- Vision Generalist Model: A Survey
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
- Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder
- Region-based Cluster Discrimination for Visual Representation Learning
- LMM-Det: Make Large Multimodal Models Excel in Object Detection
- HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
- KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
- AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
- LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
- UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
- MedPrompt: LLM-CNN Fusion with Weight Routing for Medical Image Segmentation and Classification
- Synthetic Visual Genome
- PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
- Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
- Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
- Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
- LlamaSeg: Image Segmentation via Autoregressive Mask Generation
- MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
- UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
- Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced Covering
- Vision language models have difficulty recognizing virtual objects
- Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
- Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
- Vision as Unified Multimodal Generation
- Learning Streaming Video Representation via Multitask Training
- AutoP2C: An LLM-Based Agent Framework for Code Repository Generation from Multimodal Content in Academic Papers
- TerraMind: Large-Scale Generative Multimodality for Earth Observation
- SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
- COP-GEN-Beta: Unified Generative Modelling of COPernicus Imagery Thumbnails
Related