2019/04/23 by Michelle Guo, Guo, Michelle, Albert Haque +3
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1904.10760
openalex publication_date 2019/04/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper, we address the task of spoken language understanding. We present a method for translating spoken sentences from one language into spoken sentences in another language. Given spectrogram-spectrogram pairs, our model can be trained completely from scratch to translate unseen sentences. Our method consists of a pyramidal-bidirectional recurrent network combined with a convolutional network to output sentence-level spectrograms in the target language. Empirically, our model achieves competitive performance with state-of-the-art methods on multiple languages and can generalize to unseen speakers.