2022/04/18 by Subba Reddy Oota, Oota, Subba Reddy, Jashn Arora +7 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Cognitive Computing and Networks #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Biological sciences #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Neurons and Cognition (q-bio.NC)
paper · pdf · doi:10.48550/arxiv.2204.08261
openalex publication_date 2022/04/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Enabling effective brain-computer interfaces requires understanding how the\nhuman brain encodes stimuli across modalities such as visual, language (or\ntext), etc. Brain encoding aims at constructing fMRI brain activity given a\nstimulus. There exists a plethora of neural encoding models which study brain\nencoding for single mode stimuli: visual (pretrained CNNs) or text (pretrained\nlanguage models). Few recent papers have also obtained separate visual and text\nrepresentation models and performed late-fusion using simple heuristics.\nHowever, previous work has failed to explore: (a) the effectiveness of image\nTransformer models for encoding visual stimuli, and (b) co-attentive\nmulti-modal modeling for visual and text reasoning. In this paper, we\nsystematically explore the efficacy of image Transformers (ViT, DEiT, and BEiT)\nand multi-modal Transformers (VisualBERT, LXMERT, and CLIP) for brain encoding.\nExtensive experiments on two popular datasets, BOLD5000 and Pereira, provide\nthe following insights. (1) To the best of our knowledge, we are the first to\ninvestigate the effectiveness of image and multi-modal Transformers for brain\nencoding. (2) We find that VisualBERT, a multi-modal Transformer, significantly\noutperforms previously proposed single-mode CNNs, image Transformers as well as\nother previously proposed multi-modal models, thereby establishing new\nstate-of-the-art. The supremacy of visio-linguistic models raises the question\nof whether the responses elicited in the visual regions are affected implicitly\nby linguistic processing even when passively viewing images. Future fMRI tasks\ncan verify this computational insight in an appropriate experimental setting.\n