2023/11/16 by Kateryna Chumachenko, Chumachenko, Kateryna, Alexandros Iosifidis +3
Computer Science · Psychology · #FOS: Computer and information sciences #Hand Gesture Recognition Systems #Hearing Impairment and Communication #Machine Learning (cs.LG) #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.2311.10170
openalex publication_date 2023/11/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper proposes an approach for improving performance of unimodal models with multimodal training. Our approach involves a multi-branch architecture that incorporates unimodal models with a multimodal transformer-based branch. By co-training these branches, the stronger multimodal branch can transfer its knowledge to the weaker unimodal branches through a multi-task objective, thereby improving the performance of the resulting unimodal models. We evaluate our approach on tasks of dynamic hand gesture recognition based on RGB and Depth, audiovisual emotion recognition based on speech and facial video, and audio-video-text based sentiment analysis. Our approach outperforms the conventionally trained unimodal counterparts. Interestingly, we also observe that optimization of the unimodal branches improves the multimodal branch, compared to a similar multimodal model trained from scratch.