vix.ing · top · new · best · stats · spec

Coreference Resolution for Full German Novels using Large Language Models

2026/04/21 by Agnes Hilger, Anton Ehrmanntraut · 1 voice
Computer Science · Social Sciences · #Authorship Attribution and Profiling #Computational and Text Analysis Methods #Sentiment Analysis and Opinion Mining

paper · doi:10.26083/tuda-7983

openalex publication_date 2026/04/21 · openalex created_date 2026/05/06 · openalex updated_date 2026/07/29

Abstract

The paper introduces the GerFuN dataset, consisting of five German-language novels fully annotated for character coreference, comprising a total of 450,000 tokens. Using a semi-manual pipeline, we first pre-annotated the novels using a LLM, and then manually corrected the annotations in the INCEpTION tool. The annotation guidelines, which build on existing approaches but are made more explicit and refined, are presented in the paper and released alongside the dataset. Finally, we evaluate LLMs on GerFuN, which surpass previous pipelines and exhibit near-human accuracy on prototypical cases of particular interest to literary studies.

Discussions

Related