vix.ing · top · new · best · stats · spec

Milimili. Collecting Parallel Data via Crowdsourcing

2023/07/23 by A. A. Antonov, Antonov, Alexander
Computer Science · #Authorship Attribution and Profiling #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques

paper · pdf · doi:10.48550/arxiv.2307.12282

openalex publication_date 2023/07/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We present a methodology for gathering a parallel corpus through crowdsourcing, which is more cost-effective than hiring professional translators, albeit at the expense of quality. Additionally, we have made available experimental parallel data collected for Chechen-Russian and Fula-English language pairs.

Related