Learning to Prompt for Vision-Language Models
2021/09/30 by Kaiyang Zhou, Jingkang Yang, Chen Change Loy +1 · 459 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #Topic Modeling #cs.AI #cs.CV #cs.LG
paper · pdf · doi:10.1007/s11263-022-01653-1
International Journal of Computer Vision (IJCV), 2022. Update: Adds results on the DOSCO (DOmain Shift in COntext) benchmark
openalex publication_date 2022/07/31 · arxiv created 2022/10/06 · arxiv updated 2022/10/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
Large pre-trained vision-language models like CLIP have shown great potential in learning representations that are transferable across a wide range of downstream tasks. Different from the traditional representation learning that is based mostly on discretized labels, vision-language pre-training aligns images and texts in a common feature space, which allows zero-shot transfer to a downstream task via prompting, i.e., classification weights are synthesized from natural language describing classes of interest. In this work, we show that a major challenge for deploying such models in practice is prompt engineering, which requires domain expertise and is extremely time-consuming -- one needs to spend a significant amount of time on words tuning since a slight change in wording could have a huge impact on performance. Inspired by recent advances in prompt learning research in natural language processing (NLP), we propose Context Optimization (CoOp), a simple approach specifically for adapting CLIP-like vision-language models for downstream image recognition. Concretely, CoOp models a prompt's context words with learnable vectors while the entire pre-trained parameters are kept fixed. To handle different image recognition tasks, we provide two implementations of CoOp: unified context and class-specific context. Through extensive experiments on 11 datasets, we demonstrate that CoOp requires as few as one or two shots to beat hand-crafted prompts with a decent margin and is able to gain significant improvements over prompt engineering with more shots, e.g., with 16 shots the average gain is around 15% (with the highest reaching over 45%). Despite being a learning-based approach, CoOp achieves superb domain generalization performance compared with the zero-shot model using hand-crafted prompts.
Citations
Cited by
- ABounD: Adversarial Boundary-Driven Few-Shot Learning for Multi-Class Anomaly Detection
- GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models
- AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
- Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
- ScenarioCLIP: Pretrained Transferable Visual Language Models and Action-Genome Dataset for Natural Scene Analysis
- Collaborative Learning with Multiple Foundation Models for Source-Free Domain Adaptation
- Rethinking Plant Disease Diagnosis: Bridging the Academic-Practical Gap with Vision Transformers and Zero-Shot Learning
- Exploring Weak-to-Strong Generalization for CLIP-based Classification
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
- ATAC: Augmentation-Based Test-Time Adversarial Correction for CLIP
- ReBaPL: Repulsive Bayesian Prompt Learning
- Hierarchical Semantic Tree Anchoring for CLIP-Based Class-Incremental Learning
- Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
- QwenCLIP: Boosting Medical Vision-Language Pretraining via LLM Embeddings and Prompt tuning
- Multimodal Large Language Models as Image Classifiers
- AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models
- Medical Knowledge Intervention Prompt Tuning for Medical Image Classification
- BOFA: Bridge-Layer Orthogonal Low-Rank Fusion for CLIP-Based Class-Incremental Learning
- Text-guided Weakly Supervised Framework for Dynamic Facial Expression Recognition
- GFT: Graph Feature Tuning for Efficient Point Cloud Analysis
- Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
- Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
- Leveraging Text-Driven Semantic Variation for Robust OOD Segmentation
- MoRA: Missing Modality Low-Rank Adaptation for Visual Recognition
- SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking
- GRAVER: Generative Graph Vocabularies for Robust Graph Foundation Models Fine-tuning
- Decoupling Augmentation Bias in Prompt Learning for Vision-Language Models
- Bayesian Natural Gradient Fine-Tuning of CLIP Models via Kalman Filtering
- FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
- LGCA: Enhancing Semantic Representation via Progressive Expansion
- A Retrospect to Multi-prompt Learning across Vision and Language
- Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds
- Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection
- A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language Models
- Model Inversion with Layer-Specific Modeling and Alignment for Data-Free Continual Learning
- Understanding Hardness of Vision-Language Compositionality from A Token-level Causal Lens
- Enhancing healthcare decision support through explainable AI models for risk prediction
- Teaching Sarcasm: Few-Shot Multimodal Sarcasm Detection via Distillation to a Parameter-Efficient Student
- Latent Domain Prompt Learning for Vision-Language Models
- Test-Time Adaptive Object Detection with Foundation Model
- Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
- Few-Shot Remote Sensing Image Scene Classification with CLIP and Prompt Learning
- OpenworldAUC: Towards Unified Evaluation and Optimization for Open-world Prompt Tuning
- Adaptive Dual Prompting: Hierarchical Debiasing for Fairness-aware Graph Neural Networks
- Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
- GraphTOP: Graph Topology-Oriented Prompting for Graph Neural Networks
- BiomedXPro: Prompt Optimization for Explainable Diagnosis with Biomedical Vision Language Models
- TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection
- Bi-CoG: Bi-Consistency-Guided Self-Training for Vision-Language Models
- FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
- Revisiting Logit Distributions for Reliable Out-of-Distribution Detection
- Exploring Conditions for Diffusion models in Robotic Control
- Class-Aware Prototype Learning with Negative Contrast for Test-Time Adaptation of Vision-Language Models
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
- VeFA: Vector-Based Feature Space Adaptation for Robust Model Fine-Tuning
- ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder
- FedDEAP: Adaptive Dual-Prompt Tuning for Multi-Domain Federated Learning
- Online In-Context Distillation for Low-Resource Vision Language Models
- Experience-Driven Exploration for Efficient API-Free AI Agents
- Learning an Image Editing Model without Image Editing Pairs
- Exploring Cross-Modal Flows for Few-Shot Learning
- Sample-Aware Knowledge Association and Enhancement for Open-Vocabulary Continual Learning
- DistilCLIP-EEG: Enhancing Epileptic Seizure Detection Through Multi-modal Learning and Knowledge Distillation
- Prompt-based Adaptation in Large-scale Vision Models: A Survey
- OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
- EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking
- State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification
- ΔEnergy: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization
- Collaborative Learning of Semantic-Aware Feature Learning and Label Recovery for Multi-Label Image Recognition with Incomplete Labels
- DREAM: A Benchmark Study for Deepfake REalism AssessMent
- Cooperative Pseudo Labeling for Unsupervised Federated Classification
- Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation
- Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
- D-TPT: Dimensional Entropy Maximization for Calibrating Test-Time Prompt Tuning in Vision-Language Models
- RASALoRE: Region Aware Spatial Attention with Location-based Random Embeddings for Weakly Supervised Anomaly Detection in Brain MRI Scans
- Enhancing Visual Prompting through Expanded Transformation Space and Overfitting Mitigation
- Source-Free Cross-Domain Continual Learning
- Approximate Domain Unlearning for Vision-Language Models
- Provably Robust Adaptation for Language-Empowered Foundation Models
- Few-Shot Adaptation Benchmark for Remote Sensing Vision-Language Models
- Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning
- Conditional Representation Learning for Customized Tasks
- Referring Expression Comprehension for Small Objects
- Bayesian Test-time Adaptation for Object Recognition and Detection with Vision-language Models
- Deep Generative Continual Learning using Functional LoRA: FunLoRA
- Zero-Shot Decentralized Federated Learning
- SeMoBridge: Semantic Modality Bridge for Efficient Few-Shot Adaptation of CLIP
- MAPLE: Multi-scale Attribute-enhanced Prompt Learning for Few-shot Whole Slide Image Classification
- EMO-TTA: Improving Test-Time Adaptation of Audio-Language Models for Speech Emotion Recognition
- DAM: Dual Active Learning with Multimodal Foundation Model for Source-Free Domain Adaptation
- Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric
- GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
- Understanding Catastrophic Interference: On the Identifibility of Latent Representations
- Hierarchical Representation Matching for CLIP-based Class-Incremental Learning
- Training-Free Synthetic Data Generation with Dual IP-Adapter Guidance
- SpecXNet: A Dual-Domain Convolutional Network for Robust Deepfake Detection
- A Tale of Two Experts: Cooperative Learning for Source-Free Unsupervised Domain Adaptation
- PreLoRA: Hybrid Pre-training of Vision Transformers with Full Training and Low-Rank Adapters
- Federated Domain Generalization with Domain-specific Soft Prompts Generation
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- Continual Learning with Vision-Language Models via Semantic-Geometry Preservation
- Impromptu: a framework for model-driven prompt engineering
- Bypassing Guardrails: Lessons Learned from Red Teaming ChatGPT
- Alternating Training-based Label Smoothing Enhances Prompt Generalization
- No Labels Needed: Zero-Shot Image Classification with Collaborative Self-Learning
- Knowledge Transfer from Interaction Learning
- Vision-prompting cross-modal hashing with ranking-aware margin loss
- What Makes You Unique? Attribute Prompt Composition for Object Re-Identification
- Language-in-the-Loop Culvert Inspection on the Erie Canal
- COLA: Context-aware Language-driven Test-time Adaptation
- Training-Free Label Space Alignment for Universal Domain Adaptation
- Dual-View Alignment Learning with Hierarchical-Prompt for Class-Imbalance Multi-Label Classification
- CoDoL: Conditional Domain Prompt Learning for Out-of-Distribution Generalization
- Calibration-Aware Prompt Learning for Medical Vision-Language Models
- Multimodal Representation Learning Conditioned on Semantic Relations
- MARIC: Multi-Agent Reasoning for Image Classification
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
- Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
- Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
- UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry
- Vision-Language Models for Vision Tasks: A Survey
- Contextualized Representation Learning for Effective Human-Object Interaction Detection
- Cross-Domain Attribute Alignment with CLIP: A Rehearsal-Free Approach for Class-Incremental Unsupervised Domain Adaptation
- MCL-AD: Multimodal Collaboration Learning for Zero-Shot 3D Anomaly Detection
- Image Recognition with Vision and Language Embeddings of VLMs
- Improving Personalized Search with Regularized Low-Rank Parameter Updates
- Vision-Language Semantic Aggregation Leveraging Foundation Model for Generalizable Medical Image Segmentation
- SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
- Data-Efficient Fine-Tuning of Vision-Language Models for Diagnosis of Alzheimer's Disease
- Prompt Optimization Meets Subspace Representation Learning for Few-shot Out-of-Distribution Detection
- Multi-View Slot Attention Using Paraphrased Texts for Face Anti-Spoofing
- AttriPrompt: Dynamic Prompt Composition Learning for CLIP
- AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
- CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
- Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
- SalientFusion: Context-Aware Compositional Zero-Shot Food Recognition
- Causality-guided Prompt Learning for Vision-language Models via Visual Granulation
- Modular Embedding Recomposition for Incremental Learning
- Singular Value Few-shot Adaptation of Vision-Language Models
- Resilient Multimodal Industrial Surface Defect Detection with Uncertain Sensors Availability
- Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA
- PointAD+: Learning Hierarchical Representations for Zero-shot 3D Anomaly Detection
- RS-OOD: A Vision-Language Augmented Framework for Out-of-Distribution Detection in Remote Sensing
- Beyond Human-prompting: Adaptive Prompt Tuning with Semantic Alignment for Anomaly Detection
- Learning Yourself: Class-Incremental Semantic Segmentation with Language-Inspired Bootstrapped Disentanglement
- Target-Oriented Single Domain Generalization
- Language-Aware Information Maximization for Transductive Few-Shot CLIP
- Domain Generalization in-the-Wild: Disentangling Classification from Domain-Aware Representations
- EZ-Sort: Efficient Pairwise Comparison via Zero-Shot CLIP-Based Pre-Ordering and Human-in-the-Loop Sorting
- Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping
- NiceWebRL: a Python library for human subject experiments with reinforcement learning environments
- Multimodal Prototype Alignment for Semi-supervised Pathology Image Segmentation
- Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models
- CLIPSym: Delving into Symmetry Detection with CLIP
- Calibrating Biased Distribution in VFM-derived Latent Space via Cross-Domain Geometric Consistency
- Federated Cross-Modal Style-Aware Prompt Generation
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- AdaRing: Towards Ultra-Light Vision-Language Adaptation via Cross-Layer Tensor Ring Decomposition
- Noise Matters: Optimizing Matching Noise for Diffusion Classifiers
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- SemPT: Semantic Prompt Tuning for Vision-Language Models
- Retrieval-Augmented Prompt for OOD Detection
- CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
- MOC: Meta-Optimized Classifier for Few-Shot Whole Slide Image Classification
- DSS-Prompt: Dynamic-Static Synergistic Prompting for Few-Shot Class-Incremental Learning
- Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design
- Image selective encryption analysis using mutual information in CNN based embedding space
- Hierarchical Adaptive networks with Task vectors for Test-Time Adaptation
- STEP-FACE: Sequential TExtual–Visual Prompting for multi-task face analysis
- Adaptive Cache Enhancement for Test-Time Adaptation of Vision-Language Models
- Effortless Vision-Language Model Specialization in Histopathology without Annotation
- OpenHAIV: A Framework Towards Practical Open-World Learning
- LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation
- Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning
- Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting
- Attribute Guidance With Inherent Pseudo-label For Occluded Person Re-identification
- Unified modality separation: A vision-language framework for unsupervised domain adaptation
- ETTA: Efficient Test-Time Adaptation for Vision-Language Models through Dynamic Embedding Updates
- Generalized Few-Shot Out-of-Distribution Detection
- Robust Prompt Tuning for Vision-Language Models with Mild Semantic Noise
- UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
- CLIPVehicle: A Unified Framework for Vision-based Vehicle Search
- NEARL-CLIP: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding
- Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
- CoPS: Conditional Prompt Synthesis for Zero-Shot Anomaly Detection
- VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
- Proactive Disentangled Modeling of Trigger-Object Pairings for Backdoor Defense
- Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment
- EvoVLMA: Evolutionary Vision-Language Model Adaptation
- MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection
- Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning
- ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection
- Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language Models
- Decouple before Align: Visual Disentanglement Enhances Prompt Tuning
- Zero-Shot Anomaly Detection with Dual-Branch Prompt Selection
- Anomalous Samples for Few-Shot Anomaly Detection
- Multi-Prompt Progressive Alignment for Multi-Source Unsupervised Domain Adaptation
- UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
- Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
- DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
- A Survey on Deep Multi-Task Learning in Connected Autonomous Vehicles
- AI in Agriculture: A Survey of Deep Learning Techniques for Crops, Fisheries and Livestock
- HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
- Foundation Models and Transformers for Anomaly Detection: A Survey
- MSGCoOp: Multiple Semantic-Guided Context Optimization for Few-Shot Learning
- AU-LLM: Micro-Expression Action Unit Detection via Enhanced LLM-Based Feature Fusion
- Optimizing Active Learning in Vision-Language Models via Parameter-Efficient Uncertainty Calibration
- Latte: Collaborative Test-Time Adaptation of Vision-Language Models in Federated Learning
- One Last Attention for Your Vision-Language Model
- Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions
- HAMLET-FFD: Hierarchical Adaptive Multi-modal Learning Embeddings Transformation for Face Forgery Detection
- Rethinking Few Shot CLIP Benchmarks: A Critical Analysis in the Inductive Setting
- Beyond Class Tokens: LLM-guided Dominant Property Mining for Few-shot Classification
- TAPS : Frustratingly Simple Test Time Active Learning for VLMs
- AF-CLIP: Zero-Shot Anomaly Detection via Anomaly-Focused CLIP Adaptation
- TrackAny3D: Transferring Pretrained 3D Models for Category-unified 3D Point Cloud Tracking
- Causality-aligned Prompt Learning via Diffusion-based Counterfactual Generation
- OW-CLIP: Data-Efficient Visual Supervision for Open-World Object Detection via Human-AI Collaboration
- Quantum-Informed Machine Learning for Predicting Spatiotemporal Chaos
- Knowledge Regularized Negative Feature Tuning of Vision-Language Models for Out-of-Distribution Detection
- LAVA: Language Driven Scalable and Versatile Traffic Video Analytics
- Handling Out-of-Distribution Data: A Survey
- PTCMIL: Multiple Instance Learning via Prompt Token Clustering for Whole Slide Image Analysis
- TextSAM-EUS: Text Prompt Learning for SAM to Accurately Segment Pancreatic Tumor in Endoscopic Ultrasound
- Hierarchical Cross-modal Prompt Learning for Vision-Language Models
- Exploring Scalable Unified Modeling for General Low-Level Vision
- A Conditional Probability Framework for Compositional Zero-shot Learning
- When Person Re-Identification Meets Event Camera: A Benchmark Dataset and An Attribute-guided Re-Identification Framework
- Prototype-Guided Pseudo-Labeling with Neighborhood-Aware Consistency for Unsupervised Adaptation
- FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
- Semantic-guided Fine-tuning of Foundation Model for Long-tailed Visual Recognition
- GLAD: Generalizable Tuning for Vision-Language Models
- Multimodal-Guided Dynamic Dataset Pruning for Robust and Efficient Data-Centric Learning
- MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
- ProtoConNet: Prototypical Augmentation and Alignment for Open-Set Few-Shot Image Classification
- Text-driven Multiplanar Visual Interaction for Semi-supervised Medical Image Segmentation
- MCPL: Multi-Modal Collaborative Prompt Learning for Medical Vision-Language Model
- DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval
- Merging Smarter, Generalizing Better: Enhancing Model Merging on OOD Data
- Focus on Texture: Rethinking Pre-training in Masked Autoencoders for Medical Image Classification
- MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models
- Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
- MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning
- Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapter
- Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection
- Towards Fine-Grained Adaptation of CLIP via a Self-Trained Alignment Score
- Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations
- Calibrated and Robust Foundation Models for Vision-Language and Medical Image Tasks Under Distribution Shift
- Revisiting Pool-based Prompt Learning for Few-shot Class-incremental Learning
- SynBridge: Bridging Reaction States via Discrete Flow for Bidirectional Reaction Prediction
- Synergistic Prompting for Robust Visual Recognition with Missing Modalities
- ViLU: Learning Vision-Language Uncertainties for Failure Prediction
- Beyond the Linear Separability Ceiling: Aligning Representations in VLMs
- EPIC: Efficient Prompt Interaction for Text-Image Classification
- MADPOT: Medical Anomaly Detection with CLIP Adaptation and Partial Optimal Transport
- Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EM
- Weighted Multi-Prompt Learning with Description-free Large Language Model Distillation
- DpDNet: An Dual-Prompt-Driven Network for Universal PET-CT Segmentation
- RSRefSeg 2: Decoupling Referring Remote Sensing Image Segmentation with Foundation Models
- CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
- Integrated Structural Prompt Learning for Vision-Language Models
- Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
- pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models
- FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution Detection
- Dual Modality-Aware Gated Prompt Tuning for Few-Shot Multimodal Sarcasm Detection
- CLEP-DG: Contrastive Learning for Speech Emotion Domain Generalization via Soft Prompt Tuning
- Demystifying ChatGPT: How It Masters Genre Recognition
- Dynamic Multimodal Prototype Learning in Vision-Language Models
- Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach
- Unlearning the Noisy Correspondence Makes CLIP More Robust
- Prompt Engineering Guidelines for Using Large Language Models in Requirements Engineering
- Mitigating Class-Tail Undercoverage in Medical Vision-Language Models under Clinical Shift
- Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization
- Context-Aware Academic Emotion Dataset and Benchmark
- SCING:Towards More Efficient and Robust Person Re-Identification through Selective Cross-modal Prompt Tuning
- Unleashing the Potential of All Test Samples: Mean-Shift Guided Test-Time Adaptation
- Deception Detection Meets Vision-Language Models
- A Closer Look at Conditional Prompt Tuning for Vision-Language Models
- Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
- MadCLIP: Few-shot Medical Anomaly Detection with CLIP
- Visual Textualization for Image Prompted Object Detection
- StackCLIP: Clustering-Driven Stacked Prompt in Zero-Shot Industrial Anomaly Detection
- Mamba-FETrack V2: Revisiting State Space Model for Frame-Event based Visual Object Tracking
- The Illusion of Progress? A Critical Look at Test-Time Adaptation for Vision-Language Models
- VisualPrompter: Prompt Optimization with Visual Feedback for Text-to-Image Synthesis
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-shot Semantic Segmentation
- ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment
- Concept Pinpoint Eraser for Text-to-image Diffusion Models via Residual Attention Gate
- Prompt Mechanisms in Medical Imaging: A Comprehensive Survey
- FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition
- Prompting without Panic: Attribute-aware, Zero-shot, Test-Time Calibration
- Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
- DiMPLe -- Disentangled Multi-Modal Prompt Learning: Enhancing Out-Of-Distribution Alignment with Invariant and Spurious Feature Separation
- Multimodal Prompt Alignment for Facial Expression Recognition
- Personalized Federated Learning via Dual-Prompt Optimization and Cross Fusion
- EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning
- SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes
- SABRE-FL: Selective and Accurate Backdoor Rejection for Federated Prompt Learning
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- ChordPrompt: Orchestrating Cross-Modal Prompt Synergy for Multi-Domain Incremental Learning in CLIP
- UltraAD: Fine-Grained Ultrasound Anomaly Classification via Few-Shot CLIP Adaptation
- Generalizing vision-language models to novel domains: A comprehensive survey
- CoCoA-Mix: Confusion-and-Confidence-Aware Mixture Model for Context Optimization
- Dual-Forward Path Teacher Knowledge Distillation: Bridging the Capacity Gap Between Teacher and Student
- Adaptive Multi-prompt Contrastive Network for Few-shot Out-of-distribution Detection
- Trustworthy Few-Shot Transfer of Medical VLMs through Split Conformal Prediction
- Few-Shot, Now for Real: Medical VLMs Adaptation without Balanced Sets or Validation
- GoalLadder: Incremental Goal Discovery with Vision-Language Models
- FOCoOp: Enhancing Out-of-Distribution Robustness in Federated Prompt Learning for Vision-Language Models
- Reliable Few-shot Learning under Dual Noises
- Learning to Adapt Frozen CLIP for Few-Shot Test-Time Domain Adaptation
- Quantitative Comparison of Fine-Tuning Techniques for Pretrained Latent Diffusion Models in the Generation of Unseen SAR Images
- SUSEP-Net: Simulation-Supervised and Contrastive Learning-based Deep Neural Networks for Susceptibility Source Separation
- NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models
- EKPC: Elastic Knowledge Preservation and Compensation for Class-Incremental Learning
- Preserving Clusters in Prompt Learning for Unsupervised Domain Adaptation
- PE-MA: Parameter-Efficient Co-Evolution of Multi-Agent Systems
- IQE-CLIP: Instance-aware Query Embedding for Zero-/Few-shot Anomaly Detection in Medical Domain
- Text to Image for Multi-Label Image Recognition with Joint Prompt-Adapter Learning
- AIR: Zero-shot Generative Model Adaptation with Iterative Refinement
- VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision-Language Models
- Exploring Visual Prompting: Robustness Inheritance and Beyond
- Does CLIP perceive art the same way we do?
- Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation
- Getting pwn’d by AI: Penetration Testing with Large Language Models
- Full Conformal Adaptation of Medical Vision-Language Models
- Enabling Validation for Robust Few-Shot Recognition
- Tuning the Right Foundation Models is What you Need for Partial Label Learning
- Interpretable Few-Shot Image Classification via Prototypical Concept-Guided Mixture of LoRA Experts
- Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction
- Vocabulary-free few-shot learning for Vision-Language Models
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
- Learning Unknown Spoof Prompts for Generalized Face Anti-Spoofing Using Only Real Face Images
- Learning Knowledge-based Prompts for Robust 3D Mask Presentation Attack Detection
- OV-COAST: Cost Aggregation with Optimal Transport for Open-Vocabulary Semantic Segmentation
- Multiple Stochastic Prompt Tuning for Few-shot Adaptation under Extreme Domain Shift
- ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
- Enhancing Target-unspecific Tasks through a Features Matrix
- Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models
- Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
- Bridging Weakly-Supervised Learning and VLM Distillation: Noisy Partial Label Learning for Efficient Downstream Adaptation
- ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
- Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering
- Active Learning via Vision-Language Model Adaptation with Open Data
- PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models
- Understanding Model Reprogramming for CLIP via Decoupling Visual Prompts
- Conformal Prediction for Zero-Shot Models
- Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting
- Provably Improving Generalization of Few-Shot Models with Synthetic Data
- An Empirical Study of Federated Prompt Learning for Vision Language Model
- To Trust Or Not To Trust Your Vision-Language Model's Prediction
- CLDTracker: A Comprehensive Language Description for Visual Tracking
- Domain Adaptation of Attention Heads for Zero-shot Anomaly Detection
- Frugal Incremental Generative Modeling using Variational Autoencoders
- Adapting Segment Anything Model for Power Transmission Corridor Hazard Segmentation
- Respond to Change with Constancy: Instruction-tuning with LLM for Non-I.I.D. Network Traffic Classification
- Continual Learning on CLIP via Incremental Prompt Tuning with Intrinsic Textual Anchors
- Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging
- PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba Adapter
- Generalized and Personalized Federated Learning with Black-Box Foundation Models via Orthogonal Transformations
- Recent Advances in Out-of-Distribution Detection with CLIP-Like Models: A Survey
- DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models
- NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-Identification
- Underwater Diffusion Attention Network with Contrastive Language-Image Joint Learning for Underwater Image Enhancement
- MetaWriter: Personalized Handwritten Text Recognition Using Meta-Learned Prompt Tuning
- Dual-Path Stable Soft Prompt Generation for Domain Generalization
- FDBPL: Faster Distillation-Based Prompt Learning for Region-Aware Vision-Language Models Adaptation
- Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation
- ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification
- PromptPath: Prompt-Adaptive Computational Pathways for In-Context Learning
- Few-Shot Learning from Gigapixel Images via Hierarchical Vision-Language Alignment and Modeling
- MoAPT: Mixture of Adversarial Prompt Tuning for Vision-Language Models
- Seeing the Trees for the Forest: Rethinking Weakly-Supervised Medical Visual Grounding
- Prompt Tuning Vision Language Models with Margin Regularizer for Few-Shot Learning under Distribution Shifts
- Few-Shot Adversarial Low-Rank Fine-Tuning of Vision-Language Models
- Few-Shot Concept Prompt Learning for Segmentation Foundation Models via Visual Grounding
- Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
- Unlabeled Data vs. Pre-trained Knowledge: Rethinking SSL in the Era of Large Models
- From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
- Benchmarking Unified Face Attack Detection via Hierarchical Prompt Tuning
- Uniformity First: Uniformity-aware Test-time Adaptation of Vision-language Models against Image Corruption
- From Shots to Stories: LLM-Assisted Video Editing with Unified Language Representations
- SMFusion: Semantic-Preserving Fusion of Multimodal Medical Images for Enhanced Clinical Diagnosis
- GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
- LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
- Open Set Domain Adaptation with Vision-language models via Gradient-aware Separation
- Generative AI Literacy: Twelve Defining Competencies
- Feasibility with Language Models for Open-World Compositional Zero-Shot Learning
- Generalizable Vision-Language Few-Shot Adaptation with Predictive Prompts and Negative Learning
- Generative Muscle Stimulation: Physical Assistance by Constraining Multimodal-AI with Biomechanical Knowledge
- MSCI: Addressing CLIP's Inherent Limitations for Compositional Zero-Shot Learning
- JointDistill: Adaptive Multi-Task Distillation for Joint Depth Estimation and Scene Segmentation
- MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models
- FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language Models
- Beyond General Prompts: Automated Prompt Refinement using Contrastive Class Alignment Scores for Disambiguating Objects in Vision-Language Models
- Promoting SAM for Camouflaged Object Detection via Selective Key Point-based Guidance
- DPL: Decoupled Prototype Learning for Enhancing Robustness of Vision-Language Transformers to Missing Modalities
- Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
- Beyond CLIP Generalization: Against Forward&Backward Forgetting Adapter for Continual Learning of Vision-Language Models
- Causal Prompt Calibration Guided Segment Anything Model for Open-Vocabulary Multi-Entity Segmentation
- Learn to Think: Bootstrapping LLM Reasoning Capability Through Graph Representation Learning
- Biomed-DPT: Dual Modality Prompt Tuning for Biomedical Vision-Language Models
- CacheFL: Privacy-Preserving and Efficient Federated Cache Model Fine-Tuning for Vision-Language Models
- Hierarchical Pre-Training of Vision Encoders with Large Language Model
- Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
- Mitigating Group-Level Fairness Disparities in Federated Visual Language Models
- Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
- The Unreasonable Effectiveness of Large Language-Vision Models for Source-free Video Domain Adaptation
- AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis
- Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
- FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models
- Unveiling the Underwater World: CLIP Perception Model-Guided Underwater Image Enhancement
- Token-Level Prompt Mixture with Parameter-Free Routing for Federated Domain Generalization
- AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning
- Boosting Single-domain Generalized Object Detection via Vision-Language Knowledge Interaction
- RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
- E-InMeMo: Enhanced Prompting for Visual In-Context Learning
- CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation
- On the Effectiveness of Adaptation Strategies for VLM-Based Federated Learning in Remote Sensing
- AdaBoosting Text Prompts for Vision-Language Models
- DP2FL: Dual Prompt Personalized Federated Learning in Foundation Models
- FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing
- AVadCLIP: Audio-Visual Collaboration for Robust Video Anomaly Detection
- Domain Generalization for Face Anti-spoofing via Content-aware Composite Prompt Engineering
- VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual Staining
- AffordanceSAM: Segment Anything Once More in Affordance Grounding
- M2IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering
- LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation
- CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey
- Bayesian Principles Improve Prompt Learning In Vision-Language Models
- LIFT+: Lightweight Fine-Tuning for Long-Tail Learning
- Post-pre-training for Modality Alignment in Vision-Language Foundation Models
- Sparsity Outperforms Low-Rank Projections in Few-Shot Adaptation
- Logits DeConfusion with CLIP for Few-Shot Learning
- Beyond Words: Augmenting Discriminative Richness via Diffusions in Unsupervised Prompt Learning
- Adapting Vision Foundation Models with Cascaded Semantics
- What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective
- Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
- Crane: Context-Guided Prompt Learning and Attention Refinement for Zero-Shot Anomaly Detection
- Learning Attribute-aware Representations for Few-shot Scene Text Segmentation
- DMPT: Decoupled Modality-aware Prompt Tuning for Multi-modal Object Re-identification
- R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning
- PATFinger: Prompt-Adapted Transferable Fingerprinting against Unauthorized Multimodal Dataset Usage
- FLOSS: Free Lunch in Open-vocabulary Semantic Segmentation
- FATE: A Prompt-Tuning-Based Semi-Supervised Learning Framework for Extremely Limited Labeled Data
- Efficient Prompt Tuning for Hierarchical Ingredient Recognition
- UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval
- Bayesian Cross-Modal Alignment Learning for Few-Shot Out-of-Distribution Generalization
- Semantically Encoding Activity Labels for Context-Aware Human Activity Recognition
- Learning Optimal Prompt Ensemble for Multi-source Visual Prompt Transfer
- ZIP: An Efficient Zeroth-order Prompt Tuning for Black-box Vision-Language Models
- Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
- SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data