2021/11/11 by Jan Christian Blaise Cruz, Cruz, Jan Christian Blaise, Charibeth Cheng +1 · 2 citations
Computer Science · Health Professions · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Interpreting and Communication in Healthcare #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2111.06053
openalex publication_date 2021/11/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper, we improve on existing language resources for the low-resource Filipino language in two ways. First, we outline the construction of the TLUnified dataset, a large-scale pretraining corpus that serves as an improvement over smaller existing pretraining datasets for the language in terms of scale and topic variety. Second, we pretrain new Transformer language models following the RoBERTa pretraining technique to supplant existing models trained with small corpora. Our new RoBERTa models show significant improvements over existing Filipino models in three benchmark datasets with an average gain of 4.47% test accuracy across the three classification tasks of varying difficulty.