2020/05/20 by Ngoc-Quan Pham, Thanh-Le Ha, Pham, Ngoc-Quan +14 · 2 citations
Computer Science · Engineering · #Artificial intelligence #Audio and Speech Processing (eess.AS) #Benchmark (surveying) #Computation and Language (cs.CL) #Computer science #Encoding (memory) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Pattern recognition (psychology) #Segmentation #Sentence #Sequence labeling #Sound (cs.SD) #Speech Recognition and Synthesis #Speech recognition #Transformer #cs.CL #cs.SD #eess.AS #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2005.09940
published in arXiv (Cornell University) (Cornell University) · Submitted to Interspeech 2020
arxiv created 2020/05/20 · openalex publication_date 2020/05/20 · arxiv updated 2020/05/21 · openalex created_date 2022/07/26 · openalex updated_date 2026/08/04
Transformer models are powerful sequence-to-sequence architectures that are\ncapable of directly mapping speech inputs to transcriptions or translations.\nHowever, the mechanism for modeling positions in this model was tailored for\ntext modeling, and thus is less ideal for acoustic inputs. In this work, we\nadapt the relative position encoding scheme to the Speech Transformer, where\nthe key addition is relative distance between input states in the\nself-attention network. As a result, the network can better adapt to the\nvariable distributions present in speech data. Our experiments show that our\nresulting model achieves the best recognition result on the Switchboard\nbenchmark in the non-augmentation condition, and the best published result in\nthe MuST-C speech translation benchmark. We also show that this model is able\nto better utilize synthetic data than the Transformer, and adapts better to\nvariable sentence segmentation quality for speech translation.\n