vix.ing · top · new · best · stats

Typhoon: Thai Large Language Models

2023/12/21 by Kunat Pipatanakul, Pipatanakul, Kunat, Phatrasek Jirabovonvisut +11 · 7 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Artificial intelligence #Automatic summarization #Benchmark (surveying) #Computation and Language (cs.CL) #Computer science #FOS: Computer and information sciences #Geography #Mathematics education #Meteorology #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Natural language processing #Psychology #Resource (disambiguation) #Topic Modeling #Typhoon

paper · pdf · doi:10.48550/arxiv.2312.13951

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2023/12/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Typhoon is a series of Thai large language models (LLMs) developed specifically for the Thai language. This technical report presents challenges and insights in developing Thai LLMs, including data preparation, pretraining, instruction-tuning, and evaluation. As one of the challenges of low-resource languages is the amount of pretraining data, we apply continual training to transfer existing world knowledge from a strong LLM. To evaluate the Thai knowledge encapsulated in each model from the pretraining stage, we develop ThaiExam, a benchmark based on examinations for high-school students and investment professionals in Thailand. In addition, we fine-tune Typhoon to follow Thai instructions, and we evaluate instruction-tuned models on Thai instruction datasets as well as translation, summarization, and question-answering tasks. Experimental results on a suite of Thai benchmarks show that Typhoon outperforms all open-source Thai language models, and its performance is on par with GPT-3.5 in Thai while having only 7 billion parameters and being 2.62 times more efficient in tokenizing Thai text.

Cited by

Related