vix.ing · top · new · best · stats · spec

Vision Transformer Adapters for Generalizable Multitask Learning

2023/08/23 by Deblina Bhattacharjee, Bhattacharjee, Deblina, Sabine Süsstrunk +3 · 3 citations
Computer Science · #Advanced Neural Network Applications #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications

paper · pdf · doi:10.48550/arxiv.2308.12372

openalex publication_date 2023/08/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We introduce the first multitasking vision transformer adapters that learn generalizable task affinities which can be applied to novel tasks and domains. Integrated into an off-the-shelf vision transformer backbone, our adapters can simultaneously solve multiple dense vision tasks in a parameter-efficient manner, unlike existing multitasking transformers that are parametrically expensive. In contrast to concurrent methods, we do not require retraining or fine-tuning whenever a new task or domain is added. We introduce a task-adapted attention mechanism within our adapter framework that combines gradient-based task similarities with attention-based ones. The learned task affinities generalize to the following settings: zero-shot task transfer, unsupervised domain adaptation, and generalization without fine-tuning to novel domains. We demonstrate that our approach outperforms not only the existing convolutional neural network-based multitasking methods but also the vision transformer-based ones. Our project page is at \urlhttps://ivrl.github.io/VTAGML.

Cited by

Related