vix.ing · top · new · best · stats · spec

Automatic Lexical Simplification for Turkish

2022/01/15 by Ahmet Yavuz Uluslu, Uluslu, Ahmet Yavuz
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2201.05878

openalex publication_date 2022/01/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

In this paper, we present the first automatic lexical simplification system for the Turkish language. Recent text simplification efforts rely on manually crafted simplified corpora and comprehensive NLP tools that can analyse the target text both in word and sentence levels. Turkish is a morphologically rich agglutinative language that requires unique considerations such as the proper handling of inflectional cases. Being a low-resource language in terms of available resources and industrial-strength tools, it makes the text simplification task harder to approach. We present a new text simplification pipeline based on pretrained representation model BERT together with morphological features to generate grammatically correct and semantically appropriate word-level simplifications.

Related