2026/05/28 by Abdulkadir Gokce, Abdülkadir Gökce, Badr AlKhamissi +1 · 1 voice
Computer Science · Neuroscience · Psychology · #Action Observation and Synchronization #Artificial neural network #ENCODE #Encoder #Encoding (memory) #Face Recognition and Perception #Feature (linguistics) #Gating #Modality (human–computer interaction) #Multisensory perception and integration #cs.LG
paper · pdf · open access · doi:10.48550/arxiv.2605.29850
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2026/05/28 · arxiv published 2026/05/28 · arxiv updated 2026/05/28 · openalex created_date 2026/05/30 · openalex updated_date 2026/07/28
Recent progress in task-optimized neural networks has established encoding models as a powerful tool for predicting brain responses to naturalistic stimuli, yet most existing approaches rely on unimodal representations. The emergence of omni-modal foundation models and rich multimodal neural datasets enables encoding models that jointly integrate visual, auditory, and linguistic information across subjects. We introduce MIRAGE, a brain encoding framework for predicting whole-brain fMRI responses to naturalistic audiovisual stimuli. MIRAGE achieves state-of-the-art performance via a native multimodal backbone and adaptive feature gating across layers. These representations are then combined with a transformer-based brain encoder and a subject-specific linear head over the cortical parcels. Controlled comparisons show that natively multimodal features consistently outperform post-hoc aggregation of independent unimodal features, across architectural levels and backbones. Beyond predictive accuracy, the learned attention weights are directly inspectable to interpret the modality-specific gating profile over the backbone, and each modality traces a distinct anatomical pattern across cortex. Together, these results propose adaptive layer-wise aggregation of natively multimodal features as a generalizable, interpretable, and accurate approach for whole-brain encoding.