2018/10/16 by Robert Forkel, Johann‐Mattis List, Simon J. Greenhill +7 · 7 citations
Computer Science · Social Sciences · Arts and Humanities · #Natural Language Processing Techniques #Language and cultural evolution #Digital Humanities and Scholarship
paper · pdf · doi:10.1038/sdata.2018.205
openalex publication_date 2018/10/16 · openalex created_date 2018/10/26 · openalex updated_date 2026/08/01
The amount of available digital data for the languages of the world is constantly increasing. Unfortunately, most of the digital data are provided in a large variety of formats and therefore not amenable for comparison and re-use. The Cross-Linguistic Data Formats initiative proposes new standards for two basic types of data in historical and typological language comparison (word lists, structural datasets) and a framework to incorporate more data types (e.g. parallel texts, and dictionaries). The new specification for cross-linguistic data formats comes along with a software package for validation and manipulation, a basic ontology which links to more general frameworks, and usage examples of best practices.