An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
2026/02/24 by Anna Martin-Boyle, William Humphreys, Martha Brown +2 · 1 voice
Computer Science · #cs.HC #cs.CL
paper · pdf · doi:10.1145/3772318.3791843
Abstract
Large Language Models (LLMs) are transforming scholarly tasks like search and summarization, but their reliability remains uncertain. Current evaluation metrics for testing LLM reliability are primarily automated approaches that prioritize efficiency and scalability, but lack contextual nuance and fail to reflect how scientific domain experts assess LLM outputs in practice. We developed and validated a schema for evaluating LLM errors in scholarly question-answering systems that reflects the assessment strategies of practicing scientists. In collaboration with domain experts, we identified 20 error patterns across seven categories through thematic analysis of 68 question-answer pairs. We validated this schema through contextual inquiries with 10 additional scientists, which showed not only which errors experts naturally identify but also how structured evaluation schemas can help them detect previously overlooked issues. Domain experts use systematic assessment strategies, including technical precision testing, value-based evaluation, and meta-evaluation of their own practices. We discuss implications for supporting expert evaluation of LLM outputs, including opportunities for personalized, schema-driven tools that adapt to individual evaluation patterns and expertise levels.
Citations
- A Survey of Sustainability in Large Language Models: Applications, Economics, and Challenges
- M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
- A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
- Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
- Think Together and Work Better: Combining Humans' and LLMs' Think-Aloud Outcomes for Effective Text Evaluation
- SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
- Towards Robust Evaluation: A Comprehensive Taxonomy of Datasets and Metrics for Open Domain Question Answering in the Era of Large Language Models
- SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
- Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making
- Trust in AI: Progress, Challenges, and Future Directions
- AI-Augmented Brainwriting: Investigating the use of LLMs in group ideation
- LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models
- PROXYQA: An Alternative Framework for Evaluating Long-Form Text Generation with Large Language Models
- A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions
- Evaluating the Effectiveness of Retrieval-Augmented Large Language Models in Scientific Document Reasoning
- The Troubling Emergence of Hallucination in Large Language Models -- An Extensive Definition, Quantification, and Prescriptive Remediations
- ChainForge: A Visual Toolkit for Prompt Engineering and LLM Hypothesis Testing
- ExpertQA: Expert-Curated Questions and Attributed Answers
- Generative User-Experience Research for Developing Domain-specific Natural Language Processing Applications
- AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap
- A Critical Evaluation of Evaluations for Long-form Question Answering
- Evaluating Open-Domain Question Answering in the Era of Large Language Models
- Forecasting the future of artificial intelligence with machine learning-based link prediction in an exponentially growing knowledge network
- AI and the Everything in the Whole Wide World Benchmark
- A Dataset of Information-Seeking Questions and Answers Anchored in\n Research Papers
- Dynabench: Rethinking Benchmarking in NLP
- What Will it Take to Fix Benchmarking in Natural Language Understanding?
- Hurdles to Progress in Long-form Question Answering
- Growth rates of modern science: A latent piecewise growth curve approach to model publication numbers from established and new literature databases
- Unsupervised Evaluation for Question Answering with Transformers
- Evaluation of Text Generation: A Survey
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- A Framework for Evaluation of Machine Reading Comprehension Gold Standards
- PubMedQA: A Dataset for Biomedical Research Question Answering
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Apprenticing with the customer
Discussions