2019/12/03 by Thanapapas Horsuwan, Kasidis Kanwatchara, Horsuwan, Thanapapas +5
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Natural Language Processing Techniques #Text and Document Classification Technologies #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1912.01580
openalex publication_date 2019/12/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The ever-growing volume of data of user-generated content on social media\nprovides a nearly unlimited corpus of unlabeled data even in languages where\nresources are scarce. In this paper, we demonstrate that state-of-the-art\nresults on two Thai social text categorization tasks can be realized by\npretraining a language model on a large noisy Thai social media corpus of over\n1.26 billion tokens and later fine-tuned on the downstream classification\ntasks. Due to the linguistically noisy and domain-specific nature of the\ncontent, our unique data preprocessing steps designed for Thai social media\nwere utilized to ease the training comprehension of the model. We compared four\nmodern language models: ULMFiT, ELMo with biLSTM, OpenAI GPT, and BERT. We\nsystematically compared the models across different dimensions including speed\nof pretraining and fine-tuning, perplexity, downstream classification\nbenchmarks, and performance in limited pretraining data.\n