vix.ing · top · new · best · stats

Self-supervised Contrastive Cross-Modality Representation Learning for Spoken Question Answering

2021/09/08 by Chenyu You, Nuo Chen, You, Chenyu +3 · 1 citation
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Artificial intelligence #Audio and Speech Processing (eess.AS) #Coherence (philosophical gambling strategy) #Computation and Language (cs.CL) #Computer science #Consistency (knowledge bases) #FOS: Computer and information sciences #FOS: Electrical engineering #Feature learning #Machine Learning (cs.LG) #Natural Language Processing Techniques #Natural language processing #Representation (politics) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech recognition #Topic Modeling #Utterance #cs.AI #cs.CL #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2109.03381

published in arXiv (Cornell University) (Cornell University)

arxiv created 2021/09/08 · openalex publication_date 2021/09/08 · arxiv updated 2021/09/09 · openalex created_date 2021/11/22 · openalex updated_date 2026/08/05

Abstract

Spoken question answering (SQA) requires fine-grained understanding of both spoken documents and questions for the optimal answer prediction. In this paper, we propose novel training schemes for spoken question answering with a self-supervised training stage and a contrastive representation learning stage. In the self-supervised stage, we propose three auxiliary self-supervised tasks, including utterance restoration, utterance insertion, and question discrimination, and jointly train the model to capture consistency and coherence among speech documents without any additional data or annotations. We then propose to learn noise-invariant utterance representations in a contrastive objective by adopting multiple augmentation strategies, including span deletion and span substitution. Besides, we design a Temporal-Alignment attention to semantically align the speech-text clues in the learned common space and benefit the SQA tasks. By this means, the training schemes can more effectively guide the generation model to predict more proper answers. Experimental results show that our model achieves state-of-the-art results on three SQA benchmarks.

Citations

Cited by

Related