vix.ing · top · new · best · stats

MT-Adapted Datasheets for Datasets: Template and Repository

2020/05/27 by Marta R. Costa-jussà, Marta R. Costa‐jussà, Costa-jussà, Marta R. +14
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling #cs.CL

paper · pdf · doi:10.48550/arxiv.2005.13156

arxiv created 2020/05/27 · openalex publication_date 2020/05/27 · arxiv updated 2020/05/28 · openalex created_date 2020/06/05 · openalex updated_date 2026/07/28

Abstract

In this report we are taking the standardized model proposed by Gebru et al. (2018) for documenting the popular machine translation datasets of the EuroParl (Koehn, 2005) and News-Commentary (Barrault et al., 2019). Within this documentation process, we have adapted the original datasheet to the particular case of data consumers within the Machine Translation area. We are also proposing a repository for collecting the adapted datasheets in this research area

Citations

Related