Vision-Language Models for Vision Tasks: A Survey
2023/04/03 by Jingyi Zhang, Zhang, Jingyi, Jiaxing Huang +5 · 188 citations
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning
paper · pdf · doi:10.48550/arxiv.2304.00685
Abstract
Most visual recognition studies rely heavily on crowd-labelled data in deep neural networks (DNNs) training, and they usually train a DNN for each single visual recognition task, leading to a laborious and time-consuming visual recognition paradigm. To address the two challenges, Vision-Language Models (VLMs) have been intensively investigated recently, which learns rich vision-language correlation from web-scale image-text pairs that are almost infinitely available on the Internet and enables zero-shot predictions on various visual recognition tasks with a single VLM. This paper provides a systematic review of visual language models for various visual recognition tasks, including: (1) the background that introduces the development of visual recognition paradigms; (2) the foundations of VLM that summarize the widely-adopted network architectures, pre-training objectives, and downstream tasks; (3) the widely-adopted datasets in VLM pre-training and evaluations; (4) the review and categorization of existing VLM pre-training methods, VLM transfer learning methods, and VLM knowledge distillation methods; (5) the benchmarking, analysis and discussion of the reviewed methods; (6) several research challenges and potential research directions that could be pursued in the future VLM studies for visual recognition. A project associated with this survey has been created at https://github.com/jingyi0000/VLMsurvey.
Cited by
- Visual Autoregressive Modelling for Monocular Depth Estimation
- Development of Vision-Language Model-based GNSS Spoofing Detection for Autonomous Vehicle Navigation
- CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding
- Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
- Towards Responsible and Explainable AI Agents with Consensus-Driven Reasoning
- FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
- HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
- Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
- Vision-Language Model Guided Image Restoration
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- CitySeeker: How Do VLMS Explore Embodied Urban Navigation With Implicit Human Needs?
- TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models
- Collaborative Edge-to-Server Inference for Vision-Language Models
- Are vision-language models ready to zero-shot replace supervised classification models in agriculture?
- Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
- Advancing Cache-Based Few-Shot Classification via Patch-Driven Relational Gated Graph Attention
- MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
- Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models
- COVLM-RL: Critical Object-Oriented Reasoning for Autonomous Driving Using VLM-Guided Reinforcement Learning
- SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
- A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows
- HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
- Is Generation Required for Data-Efficient Perception?
- Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery
- ASTRIDE: A Security Threat Modeling Platform for Agentic-AI Applications
- Embodied Co-Design for Rapidly Evolving Agents: Taxonomy, Frontiers, and Challenges
- I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing Models
- Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
- DepthScape: Authoring 2.5D Designs via Depth Estimation, Semantic Understanding, and Geometry Extraction
- CauSight: Learning to Supersense for Visual Causal Discovery
- IGen: Scalable Data Generation for Robot Learning from Open-World Images
- SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models
- When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI
- Artwork Interpretation with Vision Language Models: A Case Study on Emotions and Emotion Symbols
- Leveraging Textual Compositional Reasoning for Robust Change Captioning
- Text-guided Controllable Diffusion for Realistic Camouflage Images Generation
- Gender Bias in Emotion Recognition by Large Language Models
- Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
- Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation
- AVERY: Adaptive VLM Split Computing through Embodied Self-Awareness for Efficient Disaster Response Systems
- Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
- ActDistill: General Action-Guided Self-Derived Distillation for Efficient Vision-Language-Action Models
- Generative Model Predictive Control in Manufacturing Processes: A Review
- TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots
- An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
- LLaVA3: Representing 3D Scenes like a Cubist Painter to Boost 3D Scene Understanding of VLMs
- Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities
- Jailbreaking Large Vision Language Models in Intelligent Transportation Systems
- Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
- Semantic Document Derendering: SVG Reconstruction via Vision-Language Modeling
- MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications
- Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing
- Medical Knowledge Intervention Prompt Tuning for Medical Image Classification
- SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
- VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
- Feature Quality and Adaptability of Medical Foundation Models: A Comparative Evaluation for Radiographic Classification and Segmentation
- Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
- Leveraging Text-Driven Semantic Variation for Robust OOD Segmentation
- S2LM: Towards Semantic Steganography via Large Language Models
- GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-trained Robot Policy Enhancement
- When Generative Artificial Intelligence meets Extended Reality: A Systematic Review
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- Real-IAD Variety: Pushing Industrial Anomaly Detection Dataset to a Modern Era
- FedReplay: A Feature Replay Assisted Federated Transfer Learning Framework for Efficient and Privacy-Preserving Smart Agriculture
- ECVL-ROUTER: Scenario-Aware Routing for Vision-Language Models
- Chain of Time: In-Context Physical Simulation with Image Generation Models
- Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA
- Enhancing Compositional Reasoning in CLIP via Reconstruction and Alignment of Text Descriptions
- Latent Domain Prompt Learning for Vision-Language Models
- QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models
- A Survey on Efficient Vision-Language-Action Models
- Agentsway -- Software Development Methodology for AI Agents-based Teams
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
- Token-Level Inference-Time Alignment for Vision-Language Models
- Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations
- Graph4MM: Weaving Multimodal Learning with Structural Information
- RoboGPT-R1: Enhancing Robot Planning with Reinforcement Learning
- Efficient Few-Shot Learning in Remote Sensing: Fusing Vision and Vision-Language Models
- Language as a Label: Zero-Shot Multimodal Classification of Everyday Postures under Data Scarcity
- A Text-Image Fusion Method with Data Augmentation Capabilities for Referring Medical Image Segmentation
- AgentCaster: Reasoning-Guided Tornado Forecasting
- A Survey on Agentic Multimodal Large Language Models
- Learning Dynamics of VLM Finetuning
- Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- Executable Analytic Concepts as the Missing Link Between VLM Insight and Precise Manipulation
- Provably Robust Adaptation for Language-Empowered Foundation Models
- From Data to Rewards: a Bilevel Optimization Perspective on Maximum Likelihood Estimation
- MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
- CalibCLIP: Contextual Calibration of Dominant Semantics for Text-Driven Image Retrieval
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
- A Multidisciplinary Design and Optimization (MDO) Agent Driven by Large Language Models
- MonitorVLM:A Vision Language Framework for Safety Violation Detection in Mining Operations
- TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Model-Agnostic Correctness Assessment for LLM-Generated Code via Dynamic Internal Representation Selection
- VLA Model Post-Training via Action-Chunked PPO and Self Behavior Cloning
- SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
- Probing the Limits of Stylistic Alignment in Vision-Language Models
- Multimodal Arabic Captioning with Interpretable Visual Concept Integration
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis
- Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models
- MMPB: It's Time for Multi-Modal Personalization
- Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models
- An LLM-Powered Agent for Real-Time Analysis of the Vietnamese IT Job Market
- Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
- EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
- Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
- DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis
- Paper Espresso: From Paper Overload to Research Insight
- Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
- Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions
- Enhancing Speech Large Language Models through Reinforced Behavior Alignment
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout
- NaviSense: A Multimodal Assistive Mobile application for Object Retrieval by Persons with Visual Impairment
- ColorBlindnessEval: Can Vision-Language Models Pass Color Blindness Tests?
- COLA: Context-aware Language-driven Test-time Adaptation
- Multi-scale Temporal Prediction via Incremental Generation and Multi-agent Collaboration
- I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
- Orchestrate, Generate, Reflect: A VLM-Based Multi-Agent Collaboration Framework for Automated Driving Policy Learning
- Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
- ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
- Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery
- AdaThinkDrive: Adaptive Thinking via Reinforcement Learning for Autonomous Driving
- EZREAL: Enhancing Zero-Shot Outdoor Robot Navigation toward Distant Targets under Varying Visibility
- PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models
- Cross-Domain Attribute Alignment with CLIP: A Rehearsal-Free Approach for Class-Incremental Unsupervised Domain Adaptation
- The System Description of CPS Team for Track on Driving with Language of CVPR 2024 Autonomous Grand Challenge
- Towards Reliable and Interpretable Document Question Answering via VLMs
- Image Recognition with Vision and Language Embeddings of VLMs
- AWM-Fuse: Multi-Modality Image Fusion for Adverse Weather via Global and Local Text Perception
- MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
- Adapting Vision-Language Models for Neutrino Event Classification in High-Energy Physics
- Pathology-Informed Latent Diffusion Model for Anomaly Detection in Lymph Node Metastasis
- Systematic Review and Meta-analysis of AI-driven MRI Motion Artifact Detection and Correction
- Can VLMs Recall Factual Associations From Visual References?
- MoPEQ: Mixture of Mixed Precision Quantized Experts
- Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
- Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction
- Prompting with Sign Parameters for Low-resource Sign Language Instruction Generation
- Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants
- DGL-RSIS: Decoupling Global Spatial Context and Local Class Semantics for Training-Free Remote Sensing Image Segmentation
- Language-Aware Information Maximization for Transductive Few-Shot CLIP
- Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
- More Reliable Pseudo-labels, Better Performance: A Generalized Approach to Single Positive Multi-label Learning
- JVLGS: Joint Vision-Language Gas Leak Segmentation
- Linking heterogeneous microstructure informatics with expert characterization knowledge through customized and hybrid vision-language representations for industrial qualification
- Fine-Tuning Vision-Language Models for Neutrino Event Analysis in High-Energy Physics Experiments
- CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering
- DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- A Comprehensive Review of Agricultural Parcel and Boundary Delineation from Remote Sensing Images: Recent Progress and Future Perspectives
- Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
- Structured Prompting and Multi-Agent Knowledge Distillation for Traffic Video Interpretation and Risk Inference
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
- Standardization of Neuromuscular Reflex Analysis -- Role of Fine-Tuned Vision-Language Model Consortium and OpenAI gpt-oss Reasoning LLM Enabled Decision Support System
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- Deep Learning for Crack Detection: A Review of Learning Paradigms, Generalizability, and Datasets
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- Text-conditioned State Space Model For Domain-generalized Change Detection Visual Question Answering
- Re:Verse -- Can Your VLM Read a Manga?
- Towards Scalable Training for Handwritten Mathematical Expression Recognition
- Designing Object Detection Models for TinyML: Foundations, Comparative Analysis, Challenges, and Emerging Solutions
- RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
- VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
- PASG: A Closed-Loop Framework for Automated Geometric Primitive Extraction and Semantic Anchoring in Robotic Manipulation
- Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- PET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography
- OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
- MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
- Artificial Intelligence and Misinformation in Art: Can Vision Language Models Judge the Hand or the Machine Behind the Canvas?
- Context-based Motion Retrieval using Open Vocabulary Methods for Autonomous Driving
- GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- From Image Captioning to Visual Storytelling
- Enabling Few-Shot Alzheimer's Disease Diagnosis on Biomarker Data with Tabular LLMs
- On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
- A Survey on Deep Multi-Task Learning in Connected Autonomous Vehicles
- Color as the Impetus: Transforming Few-Shot Learner
Related