vix.ing · top · new · best · stats

FRACAS: A FRench Annotated Corpus of Attribution relations in newS

2023/09/19 by Ange Richard, Richard, Ange, Laura Alonzo-Canul +3 · 1 voice · 1 citation
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #cs.CL

paper · pdf · doi:10.48550/arxiv.2309.10604

openalex publication_date 2023/09/19 · arxiv published 2023/09/19 · arxiv updated 2023/09/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01

Abstract

Quotation extraction is a widely useful task both from a sociological and from a Natural Language Processing perspective. However, very little data is available to study this task in languages other than English. In this paper, we present a manually annotated corpus of 1676 newswire texts in French for quotation extraction and source attribution. We first describe the composition of our corpus and the choices that were made in selecting the data. We then detail the annotation guidelines and annotation process, as well as a few statistics about the final corpus and the obtained balance between quote types (direct, indirect and mixed, which are particularly challenging). We end by detailing our inter-annotator agreement between the 8 annotators who worked on manual labelling, which is substantially high for such a difficult linguistic phenomenon.

Cited by

Discussions

Related