vix.ing · top · new · best · stats · spec

Lexibank/Uralex: Uralex Basic Vocabulary Dataset

2018/10/12 by Kaj Syrjänen, Jyri Lehtinen, Outi Vesakoski +5 · 1 citation
Arts and Humanities · Computer Science · #Linguistics and Cultural Studies #Natural Language Processing Techniques #Linguistics and language evolution

paper · doi:10.5281/zenodo.1459402

openalex publication_date 2018/10/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/01

Abstract

The UraLex basic vocabulary dataset has its origins in the basic vocabulary cognacy dataset collected by the research initiative BEDLAN (Biological Evolution and the Diversification of Languages), funded by the Kone Foundation between 2009-2013. The data has since been revised and expanded in follow-up research projects, including SumuraSyyni (2014-2016), UraLex (2014-2016) and AikaSyyni (2017-2020). The dataset has been compiled especially for the purposes of quantitative language classification/historical linguistics, such as Bayesian Inference of phylogeny.

Cited by