2020/09/08 by Samuel Louvan, Louvan, Samuel, Bernardo Magnini +1 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2009.03695
Accepted at PACLIC 2020 - The 34th Pacific Asia Conference on Language, Information and Computation
arxiv created 2020/09/08 · openalex publication_date 2020/09/08 · arxiv updated 2020/09/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Neural-based models have achieved outstanding performance on slot filling and intent classification, when fairly large in-domain training data are available. However, as new domains are frequently added, creating sizeable data is expensive. We show that lightweight augmentation, a set of augmentation methods involving word span and sentence level operations, alleviates data scarcity problems. Our experiments on limited data settings show that lightweight augmentation yields significant performance improvement on slot filling on the ATIS and SNIPS datasets, and achieves competitive performance with respect to more complex, state-of-the-art, augmentation approaches. Furthermore, lightweight augmentation is also beneficial when combined with pre-trained LM-based models, as it improves BERT-based joint intent and slot filling models.