vix.ing · top · new · best · stats

Training Spatial-Frequency Visual Prompts and Probabilistic Clusters for\n Accurate Black-Box Transfer Learning

2024/08/15 by Wonwoo Cho, Kangyeol Kim, Cho, Wonwoo +5 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Face and Expression Recognition

paper · pdf · doi:10.48550/arxiv.2408.07944

openalex publication_date 2024/08/15 · openalex created_date 2024/09/18 · openalex updated_date 2026/07/28

Abstract

Despite the growing prevalence of black-box pre-trained models (PTMs) such as\nprediction API services, there remains a significant challenge in directly\napplying general models to real-world scenarios due to the data distribution\ngap. Considering a data deficiency and constrained computational resource\nscenario, this paper proposes a novel parameter-efficient transfer learning\nframework for vision recognition models in the black-box setting. Our framework\nincorporates two novel training techniques. First, we align the input space\n(i.e., image) of PTMs to the target data distribution by generating visual\nprompts of spatial and frequency domain. Along with the novel spatial-frequency\nhybrid visual prompter, we design a novel training technique based on\nprobabilistic clusters, which can enhance class separation in the output space\n(i.e., prediction probabilities). In experiments, our model demonstrates\nsuperior performance in a few-shot transfer learning setting across extensive\nvisual recognition datasets, surpassing state-of-the-art baselines.\nAdditionally, we show that the proposed method efficiently reduces\ncomputational costs for training and inference phases.\n

Cited by

Related