vix.ing · top · new · best · stats · spec

Improving Large-scale Language Models and Resources for Filipino

2021/11/11 by Jan Christian Blaise Cruz, Cruz, Jan Christian Blaise, Charibeth Cheng +1 · 2 citations
Computer Science · Health Professions · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Interpreting and Communication in Healthcare #Natural Language Processing Techniques #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2111.06053

openalex publication_date 2021/11/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this paper, we improve on existing language resources for the low-resource Filipino language in two ways. First, we outline the construction of the TLUnified dataset, a large-scale pretraining corpus that serves as an improvement over smaller existing pretraining datasets for the language in terms of scale and topic variety. Second, we pretrain new Transformer language models following the RoBERTa pretraining technique to supplant existing models trained with small corpora. Our new RoBERTa models show significant improvements over existing Filipino models in three benchmark datasets with an average gain of 4.47% test accuracy across the three classification tasks of varying difficulty.

Cited by

Related