2025/04/23 by Joseph M. Denning, Xiaohan Hannah Guo, Denning, Joseph M. +7 · 4 voices
Computer Science · Medicine · Neuroscience · Psychology · #Action Observation and Synchronization #Artificial Intelligence in Healthcare and Education #Cognition #Contrast (vision) #Focus (optics) #Information structure #Language model #Natural Language Processing Techniques #Neurobiology of Language and Bilingualism #Sentence #Similarity (geometry) #Thematic analysis #Thematic map #Thematic structure #Topic Modeling
paper · pdf · open access · doi:10.1162/opmi.a.365
published in Open Mind 10, 923-950 (The MIT Press)
openalex publication_date 2026/01/01 · openalex created_date 2026/07/17 · openalex updated_date 2026/07/25
Language models (LMs) are commonly criticized for not "understanding" language. However, many critiques focus on cognitive abilities that, in humans, are distinct from language processing. Here, we instead study a kind of understanding tightly linked to language: inferring "who did what to whom" (thematic roles) in a sentence. Does the central training objective of LMs-word prediction-result in sentence representations that capture thematic roles? In two experiments, we characterized sentence representations in four LMs that have been proposed as models of human language processing. The overall representational similarity of sentence pairs did not reflect whether they had the same agent/patient assignments or opposite agent/patient assignments. Furthermore, we found limited evidence that thematic role information was available in any subspace of hidden activations. However, some attention heads robustly captured thematic roles, independently of syntax. Therefore, LMs can extract thematic roles but this information influences their representations weakly.