CLIP-Adapter: Better Vision-Language Models with Feature Adapters
2021/10/09 by Peng Gao, Gao, Peng, Shijie Geng +13 · 225 citations
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Advanced Image and Video Retrieval Techniques
paper · pdf · doi:10.48550/arxiv.2110.04544
Abstract
Large-scale contrastive vision-language pre-training has shown significant progress in visual representation learning. Unlike traditional visual systems trained by a fixed set of discrete labels, a new paradigm was introduced in \citeradford2021learning to directly learn to align images with raw texts in an open-vocabulary setting. On downstream tasks, a carefully chosen text prompt is employed to make zero-shot predictions.~To avoid non-trivial prompt engineering, context optimization \citezhou2021coop has been proposed to learn continuous vectors as task-specific prompts with few-shot training examples.~In this paper, we show that there is an alternative path to achieve better vision-language models other than prompt tuning.~While prompt tuning is for the textual inputs, we propose CLIP-Adapter to conduct fine-tuning with feature adapters on either visual or language branch. Specifically, CLIP-Adapter adopts an additional bottleneck layer to learn new features and performs residual-style feature blending with the original pre-trained features.~As a consequence, CLIP-Adapter is able to outperform context optimization while maintains a simple design. Experiments and extensive ablation studies on various visual classification tasks demonstrate the effectiveness of our approach. Code is released at t https://github.com/gaopengcuhk/CLIP-Adapter.
Citations
Cited by
- Test-Time Adaptation via Dual Distillation for Videos Under Severe Distribution Shifts
- What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
- Controlling Embedding Spaces with Text-Conditioned Transformations
- CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
- Auxiliary Descriptive Knowledge for Few-Shot Adaptation of Vision-Language Model
- AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection
- In Pursuit of Pixel Supervision for Visual Pre-training
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- Improvise, Adapt, Overcome -- Telescopic Adapters for Efficient Fine-tuning of Vision Language Models in Medical Imaging
- Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
- MetaTPT: Meta Test-time Prompt Tuning for Vision-Language Models
- Advancing Cache-Based Few-Shot Classification via Patch-Driven Relational Gated Graph Attention
- Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective
- VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- Training-Free Dual Hyperbolic Adapters for Better Cross-Modal Reasoning
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks
- RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
- How (Mis)calibrated is Your Federated CLIP and What To Do About It?
- Scaling Down to Scale Up: Towards Operationally-Efficient and Deployable Clinical Models via Cross-Modal Low-Rank Adaptation for Medical Vision-Language Models
- VaMP: Variational Multi-Modal Prompt Learning for Vision-Language Models
- AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
- Generalized Out-of-Distribution Detection: A Survey
- ScenarioCLIP: Pretrained Transferable Visual Language Models and Action-Genome Dataset for Natural Scene Analysis
- Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models
- Online-PVLM: Advancing Personalized VLMs with Online Concept Learning
- Exploring Weak-to-Strong Generalization for CLIP-based Classification
- ATAC: Augmentation-Based Test-Time Adversarial Correction for CLIP
- ReBaPL: Repulsive Bayesian Prompt Learning
- EBind: a practical approach to space binding
- Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
- Medical Knowledge Intervention Prompt Tuning for Medical Image Classification
- From Classification to Cross-Modal Understanding: Leveraging Vision-Language Models for Fine-Grained Renal Pathology
- Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
- vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMs
- Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
- Federated CLIP for Resource-Efficient Heterogeneous Medical Image Classification
- Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling
- SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical Images
- SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking
- Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
- Bayesian Natural Gradient Fine-Tuning of CLIP Models via Kalman Filtering
- SegDebias: Test-Time Bias Mitigation for ViT-Based CLIP via Segmentation
- FPS: Feedforward-based Parameter Selection For Efficient Fine-Tuning
- Model Inversion with Layer-Specific Modeling and Alignment for Data-Free Continual Learning
- Teaching Sarcasm: Few-Shot Multimodal Sarcasm Detection via Distillation to a Parameter-Efficient Student
- EA3D: Online Open-World 3D Object Extraction from Streaming Videos
- Improving Visual Discriminability of CLIP for Training-Free Open-Vocabulary Semantic Segmentation
- Bi-CoG: Bi-Consistency-Guided Self-Training for Vision-Language Models
- FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
- Class-Aware Prototype Learning with Negative Contrast for Test-Time Adaptation of Vision-Language Models
- Data-Centric Lessons To Improve Speech-Language Pretraining
- VeFA: Vector-Based Feature Space Adaptation for Robust Model Fine-Tuning
- Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning
- Exploring Cross-Modal Flows for Few-Shot Learning
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
- VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
- ΔEnergy: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization
- Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation
- A Multimodal Depth-Aware Method For Embodied Reference Understanding
- Provably Robust Adaptation for Language-Empowered Foundation Models
- Few-Shot Adaptation Benchmark for Remote Sensing Vision-Language Models
- Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
- Conditional Representation Learning for Customized Tasks
- Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Classifier-Centric Adaptive Framework for Open-Vocabulary Camouflaged Object Segmentation
- Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
- GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
- Toward a Holistic Approach to Continual Model Merging
- Understanding Catastrophic Interference: On the Identifibility of Latent Representations
- 3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
- No Labels Needed: Zero-Shot Image Classification with Collaborative Self-Learning
- Knowledge Transfer from Interaction Learning
- Global Minimizers of Sigmoid Contrastive Loss
- COLA: Context-aware Language-driven Test-time Adaptation
- Informative Text-Image Alignment for Visual Affordance Learning with Foundation Models
- CoDoL: Conditional Domain Prompt Learning for Out-of-Distribution Generalization
- Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
- An Empirical Analysis of VLM-based OOD Detection: Mechanisms, Advantages, and Sensitivity
- Improving Personalized Search with Regularized Low-Rank Parameter Updates
- BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion
- Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation
- Data-Efficient Fine-Tuning of Vision-Language Models for Diagnosis of Alzheimer's Disease
- AttriPrompt: Dynamic Prompt Composition Learning for CLIP
- Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
- Singular Value Few-shot Adaptation of Vision-Language Models
- Target-Oriented Single Domain Generalization
- Language-Aware Information Maximization for Transductive Few-Shot CLIP
- Domain Generalization in-the-Wild: Disentangling Classification from Domain-Aware Representations
- Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian Alignment
- Feature-Space Planes Searcher: A Universal Domain Adaptation Framework for Interpretability and Computational Efficiency
- Glo-VLMs: Leveraging Vision-Language Models for Fine-Grained Diseased Glomerulus Classification
- Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
- SPANER: Shared Prompt Aligner for Multimodal Semantic Representation
- Preserve and Sculpt: Manifold-Aligned Fine-tuning of Vision-Language Models for Few-Shot Learning
- Cross-Domain Few-Shot Learning via Multi-View Collaborative Optimization with Vision-Language Models
- AdaRing: Towards Ultra-Light Vision-Language Adaptation via Cross-Layer Tensor Ring Decomposition
- Vision-Language Models display a strong gender bias
- MPT: Motion Prompt Tuning for Micro-Expression Recognition
- Dual-stream Network for Visual Recognition
- AME: Aligned Manifold Entropy for Robust Vision-Language Distillation
- Separation and Collaboration: Two-Level Routing Grouped Mixture-of-Experts for Multi-Domain Continual Learning
- Effortless Vision-Language Model Specialization in Histopathology without Annotation
- LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning
- Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning
- Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- Unified modality separation: A vision-language framework for unsupervised domain adaptation
- Accelerating Conditional Prompt Learning via Masked Image Modeling for Vision-Language Models
- ETTA: Efficient Test-Time Adaptation for Vision-Language Models through Dynamic Embedding Updates
- Robust Prompt Tuning for Vision-Language Models with Mild Semantic Noise
- FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding
- UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
- Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
- EgoPrompt: Prompt Learning for Egocentric Action Recognition
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
- Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual Learning
- Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language Models
- Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
- CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
- HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
- MSGCoOp: Multiple Semantic-Guided Context Optimization for Few-Shot Learning
- Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
- One Last Attention for Your Vision-Language Model
- Group Relative Augmentation for Data Efficient Action Detection
- HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection
- Rethinking Few Shot CLIP Benchmarks: A Critical Analysis in the Inductive Setting
- Regularizing Subspace Redundancy of Low-Rank Adaptation
- Beyond Class Tokens: LLM-guided Dominant Property Mining for Few-shot Classification
- TrackAny3D: Transferring Pretrained 3D Models for Category-unified 3D Point Cloud Tracking
- Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection
- DepthDark: Robust Monocular Depth Estimation for Low-Light Environments
- TextSAM-EUS: Text Prompt Learning for SAM to Accurately Segment Pancreatic Tumor in Endoscopic Ultrasound
- A Conditional Probability Framework for Compositional Zero-shot Learning
- FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
- GLAD: Generalizable Tuning for Vision-Language Models
- World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving
- Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization
- Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapter
- BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural Embeddings
- HMID-Net: An Exploration of Masked Image Modeling and Knowledge Distillation in Hyperbolic Space
- Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models
- Sparse-Dense Side-Tuner for efficient Video Temporal Grounding
- Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM
- Improving Robustness of Foundation Models in Domain Adaptation with Soup-Adapters
- Integrated Structural Prompt Learning for Vision-Language Models
- Dynamic Rank Adaptation for Vision-Language Models
- pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models
- MVL-Loc: Leveraging Vision-Language Model for Generalizable Multi-Scene Camera Relocalization
- Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach
- SCING:Towards More Efficient and Robust Person Re-Identification through Selective Cross-modal Prompt Tuning
- A Closer Look at Conditional Prompt Tuning for Vision-Language Models
- Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
- Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
- The Confidence Paradox: Can LLM Know When It's Wrong
- Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
- Continual Retinal Vision-Language Pre-training upon Incremental Imaging Modalities
- UltraAD: Fine-Grained Ultrasound Anomaly Classification via Few-Shot CLIP Adaptation
- Holmes: Towards Effective and Harmless Model Ownership Verification to Personalized Large Vision Models via Decoupling Common Features
- Generalizing vision-language models to novel domains: A comprehensive survey
- Orthogonal Projection Subspace to Aggregate Online Prior-knowledge for Continual Test-time Adaptation
- Trustworthy Few-Shot Transfer of Medical VLMs through Split Conformal Prediction
- Few-Shot, Now for Real: Medical VLMs Adaptation without Balanced Sets or Validation
- GoalLadder: Incremental Goal Discovery with Vision-Language Models
- Fine-grained Image Retrieval via Dual-Vision Adaptation
- NoLoCo: No-all-reduce Low Communication Training Method for Large Models
- Does CLIP perceive art the same way we do?
- Full Conformal Adaptation of Medical Vision-Language Models
- Enabling Validation for Robust Few-Shot Recognition
- Interpretable Few-Shot Image Classification via Prototypical Concept-Guided Mixture of LoRA Experts
- TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
- Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction
- Vocabulary-free few-shot learning for Vision-Language Models
- Multiple Stochastic Prompt Tuning for Few-shot Adaptation under Extreme Domain Shift
- Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
- Bridging Weakly-Supervised Learning and VLM Distillation: Noisy Partial Label Learning for Efficient Downstream Adaptation
- FORLA: Federated Object-centric Representation Learning with Slot Attention
- Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
- Continual-MEGA: A Large-scale Benchmark for Generalizable Continual Anomaly Detection
- PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models
- Conformal Prediction for Zero-Shot Models
- Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
- Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model
- LADA: Scalable Label-Specific CLIP Adapter for Continual Learning
- Continual Learning on CLIP via Incremental Prompt Tuning with Intrinsic Textual Anchors
- Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
- Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging
- VGLD: Visually-Guided Linguistic Disambiguation for Monocular Depth Scale Recovery
- DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models
- Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
- Dual-Path Stable Soft Prompt Generation for Domain Generalization
- Few-Shot Learning from Gigapixel Images via Hierarchical Vision-Language Alignment and Modeling
- Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
- Few-Shot Adversarial Low-Rank Fine-Tuning of Vision-Language Models
- Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
- Mamba-Adaptor: State Space Model Adaptor for Visual Recognition
- From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
- Generalizable Vision-Language Few-Shot Adaptation with Predictive Prompts and Negative Learning
- MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models
- Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis
- From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation
- Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
- Beyond CLIP Generalization: Against Forward&Backward Forgetting Adapter for Continual Learning of Vision-Language Models
- Task-Adapter++: Task-specific Adaptation with Order-aware Alignment for Few-shot Action Recognition
- Biomed-DPT: Dual Modality Prompt Tuning for Biomedical Vision-Language Models
- Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
- Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
- Interpreting Contrastive Embeddings in Specific Domains with Fuzzy Rules
- Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
- FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
- E-InMeMo: Enhanced Prompting for Visual In-Context Learning
- FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing
- LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation
- CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey
- Sparsity Outperforms Low-Rank Projections in Few-Shot Adaptation
- Logits DeConfusion with CLIP for Few-Shot Learning
- LVLMCSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning Segmentation
- FLOSS: Free Lunch in Open-vocabulary Semantic Segmentation
- RGB-Event based Pedestrian Attribute Recognition: A Benchmark Dataset and An Asymmetric RWKV Fusion Framework
- FATE: A Prompt-Tuning-Based Semi-Supervised Learning Framework for Extremely Limited Labeled Data
- UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval
- Bayesian Cross-Modal Alignment Learning for Few-Shot Out-of-Distribution Generalization
- Saliency-Motion Guided Trunk-Collateral Network for Unsupervised Video Object Segmentation
- The bottlenecks of AI: challenges for embedded and real-time research in a data-centric age
Related