PaliGemma: A versatile 3B VLM for transfer
2024/07/10 by Lucas Beyer, Beyer, Lucas, Andreas Steiner +72 · 1 voice · 166 citations
Engineering · #cs.CV #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2407.07726
Abstract
PaliGemma is an open Vision-Language Model (VLM) that is based on the SigLIP-So400m vision encoder and the Gemma-2B language model. It is trained to be a versatile and broadly knowledgeable base model that is effective to transfer. It achieves strong performance on a wide variety of open-world tasks. We evaluate PaliGemma on almost 40 diverse tasks including standard VLM benchmarks, but also more specialized tasks such as remote-sensing and segmentation.
Cited by
- Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
- A Pragmatic VLA Foundation Model
- N0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens
- The Spatial Blindspot of Vision-Language Models
- CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
- StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
- LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
- Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation
- Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
- ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
- Collaborative Edge-to-Server Inference for Vision-Language Models
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- BLURR: A Boosted Low-Resource Inference for Vision-Language-Action Models
- Benchmarking the Generality of Vision-Language-Action Models
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
- Is Generation Required for Data-Efficient Perception?
- No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
- Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
- CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation
- M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
- HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
- Jina-VLM: Small Multilingual Vision Language Model
- Hierarchical Vision Language Action Model Using Success and Failure Demonstrations
- Colon-X: Advancing Intelligent Colonoscopy toward Clinical Reasoning
- SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
- Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
- VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
- E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Continuized Discrete Diffusion
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
- Text-Guided Semantic Image Encoder
- Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
- MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
- Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
- Fara-7B: An Efficient Agentic Model for Computer Use
- Mixture of Horizons in Action Chunking
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- InternData-A1: Pioneering High-Fidelity Synthetic Data for Pre-training Generalist Policy
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- Attention Grounded Enhancement for Visual Document Retrieval
- AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
- Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque
- ViPRA: Video Prediction for Robot Actions
- How Do VLAs Effectively Inherit from VLMs?
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- VinDr-CXR-VQA: A Visual Question Answering Dataset for Explainable Chest X-Ray Analysis with Multi-Task Learning
- iFlyBot-VLA Technical Report
- RzenEmbed: Towards Comprehensive Multimodal Retrieval
- RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
- Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
- Running VLAs at Real-time Speed
- Human-in-the-loop Online Rejection Sampling for Robotic Manipulation
- π_
RL: Online RL Fine-tuning for Flow-based Vision-Language-Action Models - Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- Retrieval-Augmented Search for Large-Scale Map Collections with ColPali
- Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
- A multimodal whole-slide foundation model for pathology
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- A Survey on Efficient Vision-Language-Action Models
- Dexbotic: Open-Source Vision-Language-Action Toolbox
- Data-Centric Lessons To Improve Speech-Language Pretraining
- [De|Re]constructing VLMs' Reasoning in Counting
- GigaBrain-0: A World Model-Powered Vision-Language-Action Model
- VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- Benchmarking Multimodal Large Language Models for Face Recognition
- QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
- DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning
- RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation
- Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models
- Simple Projection Variants Improve ColBERT Performance
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- Evaluating Open-Source Vision-Language Models for Multimodal Sarcasm Detection
- More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
- Topological Alignment of Shared Vision-Language Embedding Space
- UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
- Agro-Consensus: Semantic Self-Consistency in Vision-Language Models for Crop Disease Management in Developing Countries
- Holistic Order Prediction in Natural Scenes
- Text Prompt Injection of Vision Language Models
- Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects
- PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
- A Multimodal Depth-Aware Method For Embodied Reference Understanding
- IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Benchmark It Yourself (BIY): Preparing a Dataset and Benchmarking AI Models for Scatterplot-Related Tasks
- VENTURA: Adapting Image Diffusion Models for Unified Task Conditioned Navigation
- Verifier-free Test-Time Sampling for Vision-Language-Action Models
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
- Activation Quantization of Vision Encoders Needs Prefixing Registers
- ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
- Multimodal Carotid Risk Stratification with Large Vision-Language Models: Benchmarking, Fine-Tuning, and Clinical Insights
- OTR: Synthesizing Overlay Text Dataset for Text Removal
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
- DocPruner: A Storage-Efficient Framework for Multi-Vector Visual Document Retrieval via Adaptive Patch-Level Embedding Pruning
- PhysiAgent: An Embodied Agent Framework in Physical World
- Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models
- Transferring Vision-Language-Action Models to Industry Applications: Architectures, Performance, and Challenges
- RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
- Multilingual Vision-Language Models, A Survey
- Developing Vision-Language-Action Model from Egocentric Videos
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
- Generalist Robot Manipulation beyond Action Labeled Data
- Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
- FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
- Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
- Reading the unreadable: creating a dataset of 19th century English newspapers using image-to-text language models
- Latent Action Pretraining Through World Modeling
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
- RoboManipBaselines: A Unified Framework for Imitation Learning in Robotic Manipulation across Real and Simulated Environments
- Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- ORCA: Agentic Reasoning For Hallucination and Adversarial Robustness in Vision-Language Models
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- Towards Understanding Visual Grounding in Visual Language Models
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- ViewSparsifier: Killing Redundancy in Multi-View Plant Phenotyping
- Vector embedding of multi-modal texts: a tool for discovery?
- TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
- Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
- OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
- Towards Open World Detection: A Survey
- Generalist versus Specialist Vision Foundation Models for Ocular Disease and Oculomics
- Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Galaxea Open-World Dataset and G0 Dual-System VLA Model
- Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
- QuesGenie: Intelligent Multimodal Question Generation
- Enhancing Document VQA Models via Retrieval-Augmented Generation
- DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models
- Survey of Vision-Language-Action Models for Embodied Manipulation
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
- Controlling Multimodal LLMs via Reward-guided Decoding
- ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert Users
- GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
- NavA3: Understanding Any Instruction, Navigating Anywhere, Finding Anything
- RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models
- InspectVLM: Unified in Theory, Unreliable in Practice
- villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
- On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
- CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
Discussions
Related