vix.ing · top · new · best · stats

Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)

2022/03/23 by Yu Huang, Junyang Lin, Huang, Yu +7 · 29 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG

paper · pdf · doi:10.48550/arxiv.2203.12221

41 pages, 2 figures

arxiv created 2022/03/23 · openalex publication_date 2022/03/23 · arxiv updated 2022/03/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Despite the remarkable success of deep multi-modal learning in practice, it has not been well-explained in theory. Recently, it has been observed that the best uni-modal network outperforms the jointly trained multi-modal network, which is counter-intuitive since multiple signals generally bring more information. This work provides a theoretical explanation for the emergence of such performance gap in neural networks for the prevalent joint training framework. Based on a simplified data distribution that captures the realistic property of multi-modal data, we prove that for the multi-modal late-fusion network with (smoothed) ReLU activation trained jointly by gradient descent, different modalities will compete with each other. The encoder networks will learn only a subset of modalities. We refer to this phenomenon as modality competition. The losing modalities, which fail to be discovered, are the origins where the sub-optimality of joint training comes from. Experimentally, we illustrate that modality competition matches the intrinsic behavior of late-fusion joint training.

Cited by

Related