2018/05/23 by Judith Gaspers, Gaspers, Judith, Penny Karanasou +3
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning and Algorithms #Machine Learning and Data Classification #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.1805.09119
openalex publication_date 2018/05/23 · openalex created_date 2022/10/04 · openalex updated_date 2026/07/28
This paper investigates the use of Machine Translation (MT) to bootstrap a\nNatural Language Understanding (NLU) system for a new language for the use case\nof a large-scale voice-controlled device. The goal is to decrease the cost and\ntime needed to get an annotated corpus for the new language, while still having\na large enough coverage of user requests. Different methods of filtering MT\ndata in order to keep utterances that improve NLU performance and\nlanguage-specific post-processing methods are investigated. These methods are\ntested in a large-scale NLU task with translating around 10 millions training\nutterances from English to German. The results show a large improvement for\nusing MT data over a grammar-based and over an in-house data collection\nbaseline, while reducing the manual effort greatly. Both filtering and\npost-processing approaches improve results further.\n