vix.ing · top · new · best · stats · spec

Do Large Language Models Understand Literature? Case Studies and Probing Experiments on German Poetry

2025/11/30 by Fotis Jannidis, Rabea Kleymann, Julian Schröter +1 · 1 voice
Arts and Humanities · Computer Science · Social Sciences · #Artificial Intelligence in Games #Computational and Text Analysis Methods #Digital Humanities and Scholarship

paper · doi:10.48694/jcls.4225

openalex created_date 2025/12/10 · openalex publication_date 2025/12/10 · openalex updated_date 2026/07/09

Abstract

This paper explores the capabilities of large language models (LLMs) in understanding literary texts, specifically poetry, through a series of qualitative experiments. The essay has two main thrusts. On the one hand, we perform a series of probing experiments to observe the behavior of language models when asked to perform typical tasks while engaging with two German poems, analyzing these textual aspects: meter, rhyme, assonance, lexis, phrases, syntax, figurative language, titles, and meaning. The LLMs are asked to do so on three levels of interaction — general knowledge, expert knowledge, and abstraction and transfer. The more complex the understanding-related tasks that the models are supposed to solve, the more the conditions of acceptability depend on the observer and their willingness to interpret the behavior of LLMs as rational, and the more challenging it becomes to measure the LLMs' output as correct output. We will therefore embed our experiments in a discussion of the concept of understanding that allows us to meaningfully reflect what it means to say that LLMs understand a poem. On the other hand, consequently, the study seeks to adequately capture the complications of interpretation theory that arise when one explains the behavior of an LLM as an act of understanding. Our exploratory results show that LLMs excel in analyzing semantic aspects but struggle with formal elements. Performance differences exist across textual aspects rather than complexity levels. Notably, LLMs favor established interpretations over original insights and LLMs are relatively inflexible when it comes to shifting cultural perspectives unless explicitly prompted. Thus, we show the extent to which LLMs' performance covaries more with textual aspects and the extent to which it covaries with levels of task complexity.

Discussions

Related