vix.ing · top · new · best · stats

The Danish Gigaword Project

2020/05/07 by Leon Derczynski, Leon Strømberg-Derczynski, Manuel R. Ciosici +29 · 6 citations
Arts and Humanities · Computer Science · #Artificial intelligence #Cartography #Computation and Language (cs.CL) #Computer science #Corpus linguistics #Danish #FOS: Computer and information sciences #Geography #Linguistics #Natural Language Processing Techniques #Natural language processing #Scale (ratio) #Topic Modeling #Word (group theory) #cs.CL #linguistics and terminology studies

paper · pdf · doi:10.48550/arxiv.2005.03521

published in arXiv (Cornell University) (Cornell University) · Identical to the NoDaLiDa 2021 version

openalex publication_date 2020/05/07 · arxiv created 2021/05/12 · arxiv updated 2021/05/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Danish language technology has been hindered by a lack of broad-coverage corpora at the scale modern NLP prefers. This paper describes the Danish Gigaword Corpus, the result of a focused effort to provide a diverse and freely-available one billion word corpus of Danish text. The Danish Gigaword corpus covers a wide array of time periods, domains, speakers' socio-economic status, and Danish dialects.

Cited by

Related