2020/12/25 by Alex Fedorov, Fedorov, Alex, Tristan Sylvain +17
Biochemistry, Genetics and Molecular Biology · Computer Science · #AI in cancer detection #Bioinformatics and Genomic Networks #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gene expression and cancer classification #Machine Learning (cs.LG)
paper · pdf · doi:10.48550/arxiv.2012.13623
openalex publication_date 2020/12/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Sensory input from multiple sources is crucial for robust and coherent human\nperception. Different sources contribute complementary explanatory factors.\nSimilarly, research studies often collect multimodal imaging data, each of\nwhich can provide shared and unique information. This observation motivated the\ndesign of powerful multimodal self-supervised representation-learning\nalgorithms. In this paper, we unify recent work on multimodal self-supervised\nlearning under a single framework. Observing that most self-supervised methods\noptimize similarity metrics between a set of model components, we propose a\ntaxonomy of all reasonable ways to organize this process. We first evaluate\nmodels on toy multimodal MNIST datasets and then apply them to a multimodal\nneuroimaging dataset with Alzheimer's disease patients. We find that (1)\nmultimodal contrastive learning has significant benefits over its unimodal\ncounterpart, (2) the specific composition of multiple contrastive objectives is\ncritical to performance on a downstream task, (3) maximization of the\nsimilarity between representations has a regularizing effect on a neural\nnetwork, which can sometimes lead to reduced downstream performance but still\nreveal multimodal relations. Results show that the proposed approach\noutperforms previous self-supervised encoder-decoder methods based on canonical\ncorrelation analysis (CCA) or the mixture-of-experts multimodal variational\nautoEncoder (MMVAE) on various datasets with a linear evaluation protocol.\nImportantly, we find a promising solution to uncover connections between\nmodalities through a jointly shared subspace that can help advance work in our\nsearch for neuroimaging biomarkers.\n