vix.ing · top · new · best · stats · spec

TCT: A Cross-supervised Learning Method for Multimodal Sequence Representation

2019/10/23 by Wubo Li, Wei Zou, Li, Wubo +3
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Human Pose and Action Recognition #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimodal Machine Learning Applications #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1911.05186

openalex publication_date 2019/10/23 · openalex created_date 2019/11/22 · openalex updated_date 2026/07/28

Abstract

Multimodalities provide promising performance than unimodality in most tasks. However, learning the semantic of the representations from multimodalities efficiently is extremely challenging. To tackle this, we propose the Transformer based Cross-modal Translator (TCT) to learn unimodal sequence representations by translating from other related multimodal sequences on a supervised learning method. Combined TCT with Multimodal Transformer Network (MTN), we evaluate MTN-TCT on the video-grounded dialogue which uses multimodality. The proposed method reports new state-of-the-art performance on video-grounded dialogue which indicates representations learned by TCT are more semantics compared to directly use unimodality.

Citations

Related