Mondorf, Philipp
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
2024/06/26 by Anna Bavaresco, Bavaresco, Anna, Raffaella Bernardi +37 · 5 voices · 59 citations
#cs.CL
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
2024/04/02 by Mondorf, Philipp, Plank, Barbara · 25 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
- Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
2024/02/20 by Philipp Mondorf, Barbara Plank, Mondorf, Philipp +1 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Intelligent Tutoring Systems and Adaptive Learning
- The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It
2025/02/17 by Leonardo Bertolazzi, Bertolazzi, Leonardo, Philipp Mondorf +5 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
- Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
2024/06/18 by Philipp Mondorf, Barbara Plank, Mondorf, Philipp +1 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies
- Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
2024/10/02 by Philipp Mondorf, Mondorf, Philipp, Sondre Wold +3 · 4 voices · 1 citation
#cs.LG #cs.CL
- Reason to Rote: Rethinking Memorization in Reasoning
2025/07/07 by Yupei Du, Y. Le Du, Du, Yupei +10 · 1 voice · 2 citations
Computer Science · #Multi-Agent Systems and Negotiation #cs.CL #cs.LG