2026/04/21 by Agnes Hilger, Anton Ehrmanntraut · 1 voice
Computer Science · Social Sciences · #Authorship Attribution and Profiling #Computational and Text Analysis Methods #Sentiment Analysis and Opinion Mining
paper · doi:10.26083/tuda-7983
openalex publication_date 2026/04/21 · openalex created_date 2026/05/06 · openalex updated_date 2026/07/29
The paper introduces the GerFuN dataset, consisting of five German-language novels fully annotated for character coreference, comprising a total of 450,000 tokens. Using a semi-manual pipeline, we first pre-annotated the novels using a LLM, and then manually corrected the annotations in the INCEpTION tool. The annotation guidelines, which build on existing approaches but are made more explicit and refined, are presented in the paper and released alongside the dataset. Finally, we evaluate LLMs on GerFuN, which surpass previous pipelines and exhibit near-human accuracy on prototypical cases of particular interest to literary studies.