vix.ing · top · new · best · stats · spec

Diverse Linguistic Features for Assessing Reading Difficulty of Educational Filipino Texts

2021/07/31 by Joseph Marvin Imperial, Imperial, Joseph Marvin, Ethel Ong +1
Computer Science · Psychology · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reading and Literacy Development #Second Language Acquisition and Learning #Text Readability and Simplification

paper · pdf · doi:10.48550/arxiv.2108.00241

openalex publication_date 2021/07/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In order to ensure quality and effective learning, fluency, and comprehension, the proper identification of the difficulty levels of reading materials should be observed. In this paper, we describe the development of automatic machine learning-based readability assessment models for educational Filipino texts using the most diverse set of linguistic features for the language. Results show that using a Random Forest model obtained a high performance of 62.7% in terms of accuracy, and 66.1% when using the optimal combination of feature sets consisting of traditional and syllable pattern-based predictors.

Related