vix.ing · top · new · best · stats

Adaptive Cross-Modal Few-Shot Learning

2019/02/19 by Xing Chen, Chen Xing, Xing, Chen +6 · 9 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Artificial intelligence #Cancer-related molecular mechanisms research #Computer science #Context (archaeology) #Discriminative model #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Feature (linguistics) #Feature learning #Focus (optics) #Leverage (statistics) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine learning #Margin (machine learning) #Metric (unit) #Modal #Modalities #Modality (human–computer interaction) #Multimodal Machine Learning Applications #Natural language processing #Pattern recognition (psychology) #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.1902.07104

published in arXiv (Cornell University) 32, 4847-4857 (Cornell University)

openalex publication_date 2019/02/19 · openalex created_date 2019/03/02 · arxiv created 2020/02/18 · arxiv updated 2020/02/19 · openalex updated_date 2026/08/06

Abstract

Metric-based meta-learning techniques have successfully been applied to few-shot classification problems. In this paper, we propose to leverage cross-modal information to enhance metric-based few-shot learning methods. Visual and semantic feature spaces have different structures by definition. For certain concepts, visual features might be richer and more discriminative than text ones. While for others, the inverse might be true. Moreover, when the support from visual information is limited in image classification, semantic representations (learned from unsupervised text corpora) can provide strong prior knowledge and context to help learning. Based on these two intuitions, we propose a mechanism that can adaptively combine information from both modalities according to new image categories to be learned. Through a series of experiments, we show that by this adaptive combination of the two modalities, our model outperforms current uni-modality few-shot learning methods and modality-alignment methods by a large margin on all benchmarks and few-shot scenarios tested. Experiments also show that our model can effectively adjust its focus on the two modalities. The improvement in performance is particularly large when the number of shots is very small.

Cited by

Related