A Survey on Evaluation of Large Language Models
2024/01/23 by Yupeng Chang, Xu Wang, Jindong Wang +13 · 173 citations
Computer Science · Medicine · #Artificial Intelligence in Healthcare and Education #Natural Language Processing Techniques #Topic Modeling
paper · doi:10.1145/3641289
openalex publication_date 2024/01/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/31
Abstract
Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate , where to evaluate , and how to evaluate . Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, education, natural and social sciences, agent applications, and other areas. Secondly, we answer the ‘where’ and ‘how’ questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing the performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey
Citations
- Performance evaluation of classification algorithms by k-fold and leave-one-out cross validation
- Equality of Opportunity in Supervised Learning
- Deep reinforcement learning from human preferences
- SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
- Fine-Tuning Language Models from Human Preferences
- Adversarial NLI: A New Benchmark for Natural Language Understanding
- Measuring Massive Multitask Language Understanding
- GPT-3: Its Nature, Scope, Limits, and Consequences
- Making Pre-trained Language Models Better Few-shot Learners
- Measuring Mathematical Problem Solving With the MATH Dataset
- Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking
- On the Opportunities and Risks of Foundation Models
- A General Language Assistant as a Laboratory for Alignment
- Advances In Experimental Social Psychology
- Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
- PaLM: Scaling Language Modeling with Pathways
- MRKL Systems: A modular, neuro-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning
- Dynatask: A Framework for Creating Dynamic AI Benchmark Tasks
- Training language models to follow instructions with human feedback
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
- Emergent Abilities of Large Language Models
- Language Models (Mostly) Know What They Know
- Can large language models reason about medical questions?
- Aion Framework: Dimensional Emergence of AI Consciousness, Observer-Induced Collapse, and Cosmological Portal Dynamics
- Towards a Unified Multi-Dimensional Evaluator for Text Generation
- Large Language Models Encode Clinical Knowledge
- The political ideology of conversational AI: Converging evidence on ChatGPT's pro-environmental, left-libertarian orientation
- Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?
- A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT
- Dynamic Benchmarking of Masked Language Models on Temporal Concept Drift with Multiple Views
- LLaMA: Open and Efficient Foundation Language Models
- Language Is Not All You Need: Aligning Perception with Language Models
- ChatGPT for good? On opportunities and challenges of large language models for education
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
- Accuracy and Political Bias of News Source Credibility Ratings by Large Language Models
- Toxicity in ChatGPT: Analyzing Persona-assigned Language Models
- AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- Can Large Language Models Transform Computational Social Science?
- C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
- TrueTeacher: Learning Factual Consistency Evaluation with Large Language Models
- Exploring the Upper Limits of Text-Based Collaborative Filtering Using Large Language Models: Discoveries and Insights
- LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models
- Sentiment Analysis in the Era of Large Language Models: A Reality Check
- On the Planning Abilities of Large Language Models : A Critical Investigation
- FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation
- Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models
- Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback
- On Evaluating Adversarial Robustness of Large Vision-Language Models
- Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning Performance
- A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets
- Evaluating Language Models for Mathematics through Interactions
- Evaluation of AI Chatbots for Patient-Specific EHR Questions
- Benchmarking Foundation Models with Language-Model-as-an-Examiner
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization
- M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models
- Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge Evaluation
- Investigating the Effectiveness of ChatGPT in Mathematical Reasoning and Problem Solving: Evidence from the Vietnamese National High School Graduation Examination
- CMMLU: Measuring massive multitask language understanding in Chinese
- Large language models encode clinical knowledge
- Exploiting Generative AI to Scale up Intelligent Tutoring Systems
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
- A Survey of Hallucination in Large Foundation Models
- SafetyBench: Evaluating the Safety of Large Language Models
- MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
- Can Large Language Models Understand Real-World Complex Instructions?
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Cited by
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
- Large Language Models for Software Engineering: A Systematic Literature Review
- Compositionality and Sentence Meaning: Comparing Semantic Parsing and Transformers on a Challenging Sentence Similarity Dataset
- ATLAS: Harnessing retrieval-augmented generation (RAG)
- How do large-language models respond to moral dilemmas? Insights from the defining issues test
- Research Directions in Software Supply Chain Security
- State of the Art of LLM-Enabled Interaction with Visualization
- Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
- Verifying Large Language Models' Reasoning Paths via Correlation Matrix Rank
- LLMLogAnalyzer: A Clustering-Based Log Analysis Chatbot using Large Language Models
- Psychology's Questionable Research Fundamentals (QRFs): Key problems in quantitative psychology and psychological measurement beyond Questionable Research Practices (QRPs)
- AutoStreamPipe: LLM Assisted Automatic Generation of Data Stream Processing Pipelines
- Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
- Assessing and Understanding Creativity in Large Language Models
- Efficient Utility-Preserving Machine Unlearning with Implicit Gradient Surgery
- A Principle-based Framework for the Development and Evaluation of Large Language Models for Health and Wellness
- Individualized Cognitive Simulation in Large Language Models: Evaluating Different Cognitive Representation Methods
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- From Specification to Service: Accelerating API-First Development Using Multi-Agent Systems
- HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
- OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
- CLASP: Cost-Optimized LLM-based Agentic System for Phishing Detection
- Enhancing Hotel Recommendations with AI: LLM-Based Review Summarization and Query-Driven Insights
- Investigating the Impact of Dark Patterns on LLM-Based Web Agents
- Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them
- On the use of large language models in model-driven engineering
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Evaluating LLMs for Career Guidance: Comparative Analysis of Computing Competency Recommendations Across Ten African Countries
- Will AI also replace inspectors? Investigating the potential of generative AIs in usability inspection
- Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
- Towards Automatic Evaluation and Selection of PHI De-identification Models via Multi-Agent Collaboration
- The Right to Be Remembered: Preserving Maximally Truthful Digital Memory in the Age of AI
- Vision Mamba for Permeability Prediction of Porous Media
- Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
- Rethinking Evaluation in the Era of Time Series Foundation Models: (Un)known Information Leakage Challenges
- MeTA-LoRA: Data-Efficient Multi-Task Fine-Tuning for Large Language Models
- AgentCaster: Reasoning-Guided Tornado Forecasting
- MC#: Mixture Compressor for Mixture-of-Experts Large Models
- A Layered Intuition -- Method Model with Scope Extension for LLM Reasoning
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- Failure-Driven Workflow Refinement
- Improving AGI Evaluation: A Data Science Perspective
- Evaluating LLM-Based Process Explanations under Progressive Behavioral-Input Reduction
- Active Model Selection for Large Language Models
- What Is Your Agent's GPA? A Framework for Evaluating Agent Goal-Plan-Action Alignment
- GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
- From Description to Detection: LLM based Extendable O-RAN Compliant Blind DoS Detection in 5G and Beyond
- The fragility of "cultural tendencies" in LLMs
- FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
- CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
- A Set of Quebec-French Corpus of Regional Expressions and Terms
- A Multidisciplinary Design and Optimization (MDO) Agent Driven by Large Language Models
- Benchmarking Open-Source Large Language Models for Persian in Zero-Shot and Few-Shot Learning
- COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- Toward a unified framework for data-efficient evaluation of large language models
- LLM Chemistry Estimation for Multi-LLM Recommendation
- PoseGaze-AHP: A Knowledge-Based 3D Dataset for AI-Driven Ocular and Postural Diagnosis
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
- EvoSpeak: Large Language Models for Interpretable Genetic Programming-Evolved Heuristics
- Graph-S3: Enhancing Agentic textual Graph Retrieval with Synthetic Stepwise Supervision
- Color2Struct: efficient and accurate deep-learning inverse design of structural color with controllable inference
- RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation
- QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
- Benchmarking ECG Foundational Models: A Reality Check Across Clinical Tasks
- OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment
- Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
- SCI-Verifier: Scientific Verifier with Thinking
- Chat to Chip: Large Language Model Based Design of Arbitrarily Shaped Metasurfaces
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in Large Language Model Routing
- PlantGPT: An Arabidopsis‐Based Intelligent Agent that Answers Questions about Plant Functional Genomics
- Act as an expert in psychometry. The evaluation of large language models utility in psychological tests cross-cultural adaptations
- Scalable evaluation framework for retrieval augmented generation in tobacco research using large Language models
- Theory Is All You Need: AI, Human Cognition, and Causal Reasoning
- Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs
- Harnessing large language models for coding, teaching and inclusion to empower research in ecology and evolution
- Psychometrically derived 60-question benchmarks: Substantial efficiencies and the possibility of human-AI comparisons
- Smoothing-Based Conformal Prediction for Balancing Efficiency and Interpretability
- A model of errors in transformers
- Mixture of Detectors: A Compact View of Machine-Generated Text Detection
- KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI
- An LLM-Powered Agent for Real-Time Analysis of the Vietnamese IT Job Market
- SAGE: A Realistic Benchmark for Semantic Understanding
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
- UniTransfer: Video Concept Transfer via Progressive Spatial and Timestep Decomposition
- Integrated Framework for LLM Evaluation with Answer Generation
- Governance of Generative AI
- SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed Graphs
- Advancing Thesaurus Construction With Generative AI: A Structured Approach to Synonym Identification and Scope Note Development
- Universal conceptual modeling: principles, benefits, and an agenda for conceptual modeling research
- Annotation of biological samples data to standard ontologies with support from large language models
- The interplay of learning, analytics and artificial intelligence in education: A vision for hybrid intelligence
- Who is Conscious?
- <i>ToxTempAssistant</i> : using large language models to standardise cell-based toxicological test method descriptions
- DocFetch - Towards Generating Software Documentation from Multiple Software Artifacts
- Algebraic Approach to Ridge-Regularized Mean Squared Error Minimization in Minimal ReLU Neural Network
- Political DEBATE: Efficient Zero-shot and Few-shot Classifiers for Political Text
- AECBench: A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field
- Agentic AutoSurvey: Let LLMs Survey LLMs
- Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
- Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
- Bounded PCTL Model Checking of Large Language Model Outputs
- MEF: A Systematic Evaluation Framework for Text-to-Image Models
- A Knowledge Graph-based Retrieval-Augmented Generation Framework for Algorithm Selection in the Facility Layout Problem
- SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
- SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
- How ChatGPT and Gemini View the Elements of Communication Competence of Large Language Models: A Pilot Study
- Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
- Controlled Yet Natural: A Hybrid BDI-LLM Conversational Agent for Child Helpline Training
- Evaluating LLM Generated Detection Rules in Cybersecurity
- Improving Deep Tabular Learning
- Defining and Monitoring Complex Robot Activities via LLMs and Symbolic Reasoning
- Simulating a Bias Mitigation Scenario in Large Language Models
- DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models
- An Empirical Analysis of VLM-based OOD Detection: Mechanisms, Advantages, and Sensitivity
- When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective
- Multi Anatomy X-Ray Foundation Model
- Towards Deeper Understanding of Natural User Interactions in Virtual Reality Based Assembly Tasks
- Evaluating undergraduate mathematics examinations in the era of generative AI: a curriculum-level case study
- Large Language Models for Security Operations Centers: A Comprehensive Survey
- MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
- Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
- QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
- Adapting Vision-Language Models for Neutrino Event Classification in High-Energy Physics
- Investigating Student Interaction Patterns with Large Language Model-Powered Course Assistants in Computer Science Courses
- No-Knowledge Alarms for Misaligned LLMs-as-Judges
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
- Stack Overflow Is Not Dead Yet: Crowd Answers Still Matter
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- Artificial intelligence for representing and characterizing quantum systems
- AFD-SLU: Adaptive Feature Distillation for Spoken Language Understanding
- A Foundation Model for Chest X-ray Interpretation with Grounded Reasoning via Online Reinforcement Learning
- Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning
- LLM-GUARD: Large Language Model-Based Detection and Repair of Bugs and Security Vulnerabilities in C++ and Python
- Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
- Txt2Sce: Scenario Generation for Autonomous Driving System Testing Based on Textual Reports
- Federated Foundation Models in Harsh Wireless Environments: Prospects, Challenges, and Future Directions
- The Fools are Certain; the Wise are Doubtful: Exploring LLM Confidence in Code Completion
- GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
- ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
- Assessing student perceptions and use of instructor versus <scp>AI</scp> ‐generated feedback
- Fine-Tuning Vision-Language Models for Neutrino Event Analysis in High-Energy Physics Experiments
- MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation
- DELIVER: A System for LLM-Guided Coordinated Multi-Robot Pickup and Delivery using Voronoi-Based Relay Planning
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- PerFairX: Is There a Balance Between Fairness and Personality in Large Language Model Recommendations?
- The Promise of Large Language Models in Digital Health: Evidence from Sentiment Analysis in Online Health Communities
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
- Clean Code, Better Models: Enhancing LLM Performance with Smell-Cleaned Dataset
- WebGeoInfer: A Structure-Free and Multi-Stage Framework for Geolocation Inference of Devices Exposing Information
- Advancing Autonomous Incident Response: Leveraging LLMs and Cyber Threat Intelligence
- Exploring the Potential of Large Language Models in Fine-Grained Review Comment Classification
- ReqInOne: A Large Language Model-Based Agent for Software Requirements Specification Generation
- CS-Agent: LLM-based Community Search via Dual-agent Collaboration
- SinLlama -- A Large Language Model for Sinhala
- LPGNet: A Lightweight Network with Parallel Attention and Gated Fusion for Multimodal Emotion Recognition
- ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
- Cracking the Chronic Pain code: A scoping review of Artificial Intelligence in Chronic Pain research
- The Epistemic Impact of Large Language Models on Policymaking
- SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
- RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented Generation
- Fast, Convex and Conditioned Network for Multi-Fidelity Vectors and Stiff Univariate Differential Equations
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
- Hierarchical Text Classification Using Black Box Large Language Models
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
- Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
- Win-k: Improved Membership Inference Attacks on Small Language Models
- DAMR: Efficient and Adaptive Context-Aware Knowledge Graph Question Answering with LLM-Guided MCTS
- Objective Metrics for Evaluating Large Language Models Using External Data Sources
- Comparison of Large Language Models for Deployment Requirements
- On LLM-Assisted Generation of Smart Contracts from Business Processes
Related