An Introduction to Vision-Language Modeling
2024/05/27 by Florian Bordes, Bordes, Florian, Richard Yuanzhe Pang +80 · 4 voices · 72 citations
Computer Science · Social Sciences · #FOS: Computer and information sciences #Geographic Information Systems Studies #Machine Learning (cs.LG) #cs.LG
paper · pdf · doi:10.48550/arxiv.2405.17247
openalex publication_date 2024/05/27 · arxiv published 2024/05/27 · arxiv updated 2024/05/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Following the recent popularity of Large Language Models (LLMs), several attempts have been made to extend them to the visual domain. From having a visual assistant that could guide us through unfamiliar environments to generative models that produce images using only a high-level text description, the vision-language model (VLM) applications will significantly impact our relationship with technology. However, there are many challenges that need to be addressed to improve the reliability of those models. While language is discrete, vision evolves in a much higher dimensional space in which concepts cannot always be easily discretized. To better understand the mechanics behind mapping vision to language, we present this introduction to VLMs which we hope will help anyone who would like to enter the field. First, we introduce what VLMs are, how they work, and how to train them. Then, we present and discuss approaches to evaluate VLMs. Although this work primarily focuses on mapping images to language, we also discuss extending VLMs to videos.
Cited by
- Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
- In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- KathDB: Explainable Multimodal Database Management System with Human-AI Collaboration
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- Multilingual VLM Training: Adapting an English-Trained VLM to French
- GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
- HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
- Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
- Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
- Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
- Generative Model Predictive Control in Manufacturing Processes: A Review
- Zero-Shot Open-Vocabulary Human Motion Grounding with Test-Time Training
- Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
- Dataset Safety in Autonomous Driving: Requirements, Risks, and Assurance
- ZeroShotOpt: Towards Zero-Shot Pretrained Models for Efficient Black-Box Optimization
- Generalizable and scalable protein stability prediction with rewired protein generative models
- Fine-Tuning MedGemma for Clinical Captioning to Enhance Multimodal RAG over Malaysia CPGs
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- Topological Alignment of Shared Vision-Language Embedding Space
- Fundamentals of Building Autonomous LLM Agents
- RAVEN: Realtime Accessibility in Virtual ENvironments for Blind and Low-Vision People
- Reinforced Embodied Planning with Verifiable Reward for Real-World Robotic Manipulation
- GroundSight: Augmenting Vision-Language Models with Grounding Information and De-hallucination
- VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
- From Bias to Balance: Exploring and Mitigating Spatial Bias in LVLMs
- An LLM-Powered Agent for Real-Time Analysis of the Vietnamese IT Job Market
- IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
- Randomized Smoothing Meets Vision-Language Models
- Participatory AI: A Scandinavian Approach to Human-Centered AI
- Contrastive Learning with Enhanced Abstract Representations using Grouped Loss of Abstract Semantic Supervision
- AI Agents for Web Testing: A Case Study in the Wild
- AI Compute Architecture and Evolution Trends
- Prune2Drive: A Plug-and-Play Framework for Accelerating Vision-Language Models in Autonomous Driving
- M3PO: Multimodal-Model-Guided Preference Optimization for Visual Instruction Following
- Controlling Multimodal LLMs via Reward-guided Decoding
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
- Zero-shot Shape Classification of Nanoparticles in SEM Images using Vision Foundation Models
- Mining Contextualized Visual Associations from Images for Creativity Understanding
- Advancing Vision-based Human Action Recognition: Exploring Vision-Language CLIP Model for Generalisation in Domain-Independent Tasks
- Towards channel foundation models (CFMs): Motivations, methodologies and opportunities
- VIBE: Can a VLM Read the Room?
- Counterfactual Visual Explanation via Causally-Guided Adversarial Steering
- Facial Emotion Learning with Text-Guided Multiview Fusion via Vision-Language Model for 3D/4D Facial Expression Recognition
- Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
- GoalLadder: Incremental Goal Discovery with Vision-Language Models
- Using Vision Language Models to Detect Students' Academic Emotion through Facial Expressions
- Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges
- DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
- Beyond Invisibility: Learning Robust Visible Watermarks for Stronger Copyright Protection
- Can Vision Transformers with ResNet's Global Features Fairly Authenticate Demographic Faces?
- Learning Sparsity for Effective and Efficient Music Performance Question Answering
- Self-Supervised Multi-View Representation Learning using Vision-Language Model for 3D/4D Facial Expression Recognition
- Circuit Stability Characterizes Language Model Generalization
- ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
- Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- Evaluating VLMs on Multimodal Aristotelian Persuasion Tasks
- Domain Adaptation of VLM for Soccer Video Understanding
- Content Generation Models in Computational Pathology: A Comprehensive Survey on Methods, Applications, and Challenges
- Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit Hypersphere
- Unsupervised Multiview Contrastive Language-Image Joint Learning with Pseudo-Labeled Prompts Via Vision-Language Model for 3D/4D Facial Expression Recognition
- VLM-KG: Multimodal Radiology Knowledge Graph Generation
- Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
- OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding
- Reimagining Urban Science: Scaling Causal Inference with Large Language Models
- TerraMind: Large-Scale Generative Multimodality for Earth Observation
- Using Vision Language Models for Safety Hazard Identification in Construction
- Are Vision-Language Models Ready for Dietary Assessment? Exploring the Next Frontier in AI-Powered Food Image Recognition
- Towards deployment-centric multimodal AI beyond vision and language
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
Discussions
Related