2025/07/13 by Daniela Kazakouskaya, Kazakouskaya, Daniela, Timothee Mickus +3
Arts and Humanities · #Computation and Language (cs.CL) #FOS: Computer and information sciences #linguistics and terminology studies
paper · pdf · doi:10.48550/arxiv.2507.09536
openalex publication_date 2025/07/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Definition modeling, the task of generating new definitions for words in context, holds great prospect as a means to assist the work of lexicographers in documenting a broader variety of lects and languages, yet much remains to be done in order to assess how we can leverage pre-existing models for as-of-yet unsupported languages. In this work, we focus on adapting existing models to Belarusian, for which we propose a novel dataset of 43,150 definitions. Our experiments demonstrate that adapting a definition modeling systems requires minimal amounts of data, but that there currently are gaps in what automatic metrics do capture.