vix.ing · top · new · best · stats · spec

Towards Machine Comprehension of Spoken Content: Initial TOEFL Listening\n Comprehension Test by Machine

2016/08/23 by Bo-Hsiang Tseng, Sheng-syun Shen, Tseng, Bo-Hsiang +5
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Neural Networks and Applications

paper · pdf · doi:10.48550/arxiv.1608.06378

openalex publication_date 2016/08/23 · openalex created_date 2022/10/02 · openalex updated_date 2026/07/28

Abstract

Multimedia or spoken content presents more attractive information than plain\ntext content, but it's more difficult to display on a screen and be selected by\na user. As a result, accessing large collections of the former is much more\ndifficult and time-consuming than the latter for humans. It's highly attractive\nto develop a machine which can automatically understand spoken content and\nsummarize the key information for humans to browse over. In this endeavor, we\npropose a new task of machine comprehension of spoken content. We define the\ninitial goal as the listening comprehension test of TOEFL, a challenging\nacademic English examination for English learners whose native language is not\nEnglish. We further propose an Attention-based Multi-hop Recurrent Neural\nNetwork (AMRNN) architecture for this task, achieving encouraging results in\nthe initial tests. Initial results also have shown that word-level attention is\nprobably more robust than sentence-level attention for this task with ASR\nerrors.\n

Related