2017/01/10 by Abhinav Thanda, Thanda, Abhinav, Shankar M. Venkatesan +2 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #cs.AI #cs.CL #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.1701.02477
arxiv created 2017/01/10 · openalex publication_date 2017/01/10 · arxiv updated 2017/01/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Multi-task learning (MTL) involves the simultaneous training of two or more related tasks over shared representations. In this work, we apply MTL to audio-visual automatic speech recognition(AV-ASR). Our primary task is to learn a mapping between audio-visual fused features and frame labels obtained from acoustic GMM/HMM model. This is combined with an auxiliary task which maps visual features to frame labels obtained from a separate visual GMM/HMM model. The MTL model is tested at various levels of babble noise and the results are compared with a base-line hybrid DNN-HMM AV-ASR model. Our results indicate that MTL is especially useful at higher level of noise. Compared to base-line, upto 7% relative improvement in WER is reported at -3 SNR dB