2021/01/26 by Yi Zhu, Ehsan Shareghi, Zhu, Yi +7
Computer Science · Engineering · #Artificial intelligence #Artificial neural network #Bridge (graph theory) #Computation and Language (cs.CL) #Computer science #Deep learning #Economic shortage #Engineering #FOS: Computer and information sciences #Generative grammar #Generative model #Handwritten Text Recognition Techniques #Linguistics #Machine learning #Natural Language Processing Techniques #Natural language processing #Pipeline (software) #Supervised learning #Task (project management) #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2101.10717
published in arXiv (Cornell University) (Cornell University) · EACL 2021
arxiv created 2021/01/26 · openalex publication_date 2021/01/26 · arxiv updated 2021/01/27 · openalex created_date 2021/02/01 · openalex updated_date 2026/07/28
Semi-supervised learning through deep generative models and multi-lingual pretraining techniques have orchestrated tremendous success across different areas of NLP. Nonetheless, their development has happened in isolation, while the combination of both could potentially be effective for tackling task-specific labelled data shortage. To bridge this gap, we combine semi-supervised deep generative models and multi-lingual pretraining to form a pipeline for document classification task. Compared to strong supervised learning baselines, our semi-supervised classification framework is highly competitive and outperforms the state-of-the-art counterparts in low-resource settings across several languages.