2021/08/13 by Jing Zhou, Zhou, Jing, Yanan Zheng +7 · 2 citations
Computer Science · Engineering · #Computation and Language (cs.CL) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Geophysical Methods and Applications #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2108.06332
openalex publication_date 2021/08/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Most previous methods for text data augmentation are limited to simple tasks and weak baselines. We explore data augmentation on hard tasks (i.e., few-shot natural language understanding) and strong baselines (i.e., pretrained models with over one billion parameters). Under this setting, we reproduced a large number of previous augmentation methods and found that these methods bring marginal gains at best and sometimes degrade the performance much. To address this challenge, we propose a novel data augmentation method FlipDA that jointly uses a generative model and a classifier to generate label-flipped data. Central to the idea of FlipDA is the discovery that generating label-flipped data is more crucial to the performance than generating label-preserved data. Experiments show that FlipDA achieves a good tradeoff between effectiveness and robustness -- it substantially improves many tasks while not negatively affecting the others.