2024/06/01 by Zhi Zhou, Ming Yang, Zhou, Zhi +7 · 2 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #Control Systems and Identification #FOS: Computer and information sciences #Fault Detection and Control Systems #Machine Learning (cs.LG) #VLSI and Analog Circuit Testing
paper · pdf · doi:10.48550/arxiv.2406.00345
openalex publication_date 2024/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Vision-language models (VLMs), such as CLIP, have demonstrated impressive zero-shot capabilities for various downstream tasks. Their performance can be further enhanced through few-shot prompt tuning methods. However, current studies evaluate the performance of learned prompts separately on base and new classes. This evaluation lacks practicality for real-world applications since downstream tasks cannot determine whether the data belongs to base or new classes in advance. In this paper, we explore a problem setting called Open-world Prompt Tuning (OPT), which involves tuning prompts on base classes and evaluating on a combination of base and new classes. By introducing Decomposed Prompt Tuning framework (DePT), we theoretically demonstrate that OPT can be solved by incorporating out-of-distribution detection into prompt tuning, thereby enhancing the base-to-new discriminability. Based on DePT, we present a novel prompt tuning approach, namely, Decomposed Context Optimization (DeCoOp), which introduces new-class detectors and sub-classifiers to further enhance the base-class and new-class discriminability. Experimental results on 11 benchmark datasets validate the effectiveness of DePT and demonstrate that DeCoOp outperforms current state-of-the-art methods, providing a significant 2% average accuracy improvement.