A Survey on Evaluation of Large Language Models
2023/07/06 by Yupeng Chang, Xu Wang, Chang, Yupeng +29 · 220 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2307.03109
openalex publication_date 2023/07/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their evaluation becomes increasingly critical, not only at the task level, but also at the society level for better understanding of their potential risks. Over the past years, significant efforts have been made to examine LLMs from various perspectives. This paper presents a comprehensive review of these evaluation methods for LLMs, focusing on three key dimensions: what to evaluate, where to evaluate, and how to evaluate. Firstly, we provide an overview from the perspective of evaluation tasks, encompassing general natural language processing tasks, reasoning, medical usage, ethics, educations, natural and social sciences, agent applications, and other areas. Secondly, we answer the `where' and `how' questions by diving into the evaluation methods and benchmarks, which serve as crucial components in assessing performance of LLMs. Then, we summarize the success and failure cases of LLMs in different tasks. Finally, we shed light on several future challenges that lie ahead in LLMs evaluation. Our aim is to offer invaluable insights to researchers in the realm of LLMs evaluation, thereby aiding the development of more proficient LLMs. Our key point is that evaluation should be treated as an essential discipline to better assist the development of LLMs. We consistently maintain the related open-source materials at: https://github.com/MLGroupJLU/LLM-eval-survey.
Cited by
- Graph Neural Networks with Transformer Fusion of Brain Connectivity Dynamics and Tabular Data for Forecasting Future Tobacco Use
- With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
- Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Agent-based simulation of online social networks and disinformation
- Training-Driven Representational Geometry Modularization Predicts Brain Alignment in Language Models
- Self-attention vector output similarities reveal how machines pay attention
- SGCR: A Specification-Grounded Framework for Trustworthy LLM Code Review
- IntentMiner: Intent Inversion Attack via Tool Call Analysis in the Model Context Protocol
- TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
- Polypersona: Persona-Grounded LLM for Synthetic Survey Responses
- Sharing State Between Prompts and Programs
- SPARS: A Reinforcement Learning-Enabled Simulator for Power Management in HPC Job Scheduling
- Behavior and Representation in Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection
- Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
- Memoria: A Scalable Agentic Memory Framework for Personalized Conversational AI
- Human-Inspired Learning for Large Language Models via Obvious Record and Maximum-Entropy Method Discovery
- Encoder-Free Knowledge-Graph Reasoning with LLMs via Hyperdimensional Path Retrieval
- Llama-based source code vulnerability detection: Prompt engineering vs Fine tuning
- Evolutionary perspective of large language models on shaping research insights into healthcare disparities
- AutoICE: Automatically Synthesizing Verifiable C Code via LLM-driven Evolution
- CKG-LLM: LLM-Assisted Detection of Smart Contract Access Control Vulnerabilities Based on Knowledge Graphs
- Uncovering Competency Gaps in Large Language Models and Their Benchmarks
- SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code
- Are LLMs Truly Multilingual? Exploring Zero-Shot Multilingual Capability of LLMs for Information Retrieval: An Italian Healthcare Use Case
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- When Do Symbolic Solvers Enhance Reasoning in Large Language Models?
- Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
- Decentralized Multi-Agent System with Trust-Aware Communication
- LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
- Mitigating Hallucinations in Zero-Shot Scientific Summarisation: A Pilot Study
- Financial Instruction Following Evaluation (FIFE)
- Hierarchical Molecular Language Models (HMLMs)
- Invisible Hands: Gray-Box Bit Flip Attack for Steering LLMs Without Knowledge of Gradients, Data, and Weights
- CacheTrap: Injecting Trojans in LLMs without Leaving any Traces in Inputs or Weights
- Failure Modes in LLM Systems: A System-Level Taxonomy for Reliable AI Applications
- Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
- Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
- Build AI Assistants using Large Language Models and Agents to Enhance the Engineering Education of Biomechanics
- WiCo-PG: Wireless Channel Foundation Model for Pathloss Map Generation via Synesthesia of Machines
- ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning
- VLMs Guided Interpretable Decision Making for Autonomous Driving
- Translation Entropy: A Statistical Framework for Evaluating Translation Systems
- AI Agent-Driven Framework for Automated Product Knowledge Graph Construction in E-Commerce
- Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
- MACEval: A Multi-Agent Continual Evaluation Network for Large Models
- Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
- Unsteady Metrics and Benchmarking Cultures of AI Model Builders
- Evaluating Language Model Applications for Identifying Solution-Related Content in Issue Report Discussions
- LLM For Loop Invariant Generation and Fixing: How Far Are We?
- A Multi-Agent System for Semantic Mapping of Relational Data to Knowledge Graphs
- LPFQA: A Long-Tail Professional Forum-based Benchmark for LLM Evaluation
- An LLM-based Framework for Human-Swarm Teaming Cognition in Disaster Search and Rescue
- AI as We Describe It: How Large Language Models and Their Applications in Health are Represented Across Channels of Public Discourse
- When Generative Artificial Intelligence meets Extended Reality: A Systematic Review
- POLIS-Bench: Towards Multi-Dimensional Evaluation of LLMs for Bilingual Policy Tasks in Governmental Scenarios
- The Collaboration Gap
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
- Metamorphic Testing of Large Language Models for Natural Language Processing
- Evaluating Cultural Knowledge Processing in Large Language Models: A Cognitive Benchmarking Framework Integrating Retrieval-Augmented Generation
- A Hierarchical Imprecise Probability Approach to Reliability Assessment of Large Language Models
- EL-MIA: Quantifying Membership Inference Risks of Sensitive Entities in LLMs
- Chain of Time: In-Context Physical Simulation with Image Generation Models
- AutoSurvey2: Empowering Researchers with Next Level Automated Literature Surveys
- TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics Reasoning
- State of the Art of LLM-Enabled Interaction with Visualization
- ATLAS: Harnessing retrieval-augmented generation (RAG)
- Opt4GPTQ: Co-Optimizing Memory and Computation for 4-bit GPTQ Quantized LLM Inference on Heterogeneous Platforms
- Verifying Large Language Models' Reasoning Paths via Correlation Matrix Rank
- LLMLogAnalyzer: A Clustering-Based Log Analysis Chatbot using Large Language Models
- AutoStreamPipe: LLM Assisted Automatic Generation of Data Stream Processing Pipelines
- Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
- Efficient Utility-Preserving Machine Unlearning with Implicit Gradient Surgery
- A Principle-based Framework for the Development and Evaluation of Large Language Models for Health and Wellness
- Individualized Cognitive Simulation in Large Language Models: Evaluating Different Cognitive Representation Methods
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- From Specification to Service: Accelerating API-First Development Using Multi-Agent Systems
- HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
- OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
- CLASP: Cost-Optimized LLM-based Agentic System for Phishing Detection
- Enhancing Hotel Recommendations with AI: LLM-Based Review Summarization and Query-Driven Insights
- Investigating the Impact of Dark Patterns on LLM-Based Web Agents
- Presenting Large Language Models as Companions Affects What Mental Capacities People Attribute to Them
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- Evaluating LLMs for Career Guidance: Comparative Analysis of Computing Competency Recommendations Across Ten African Countries
- Will AI also replace inspectors? Investigating the potential of generative AIs in usability inspection
- Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
- Towards Automatic Evaluation and Selection of PHI De-identification Models via Multi-Agent Collaboration
- The Right to Be Remembered: Preserving Maximally Truthful Digital Memory in the Age of AI
- Vision Mamba for Permeability Prediction of Porous Media
- Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
- Rethinking Evaluation in the Era of Time Series Foundation Models: (Un)known Information Leakage Challenges
- MeTA-LoRA: Data-Efficient Multi-Task Fine-Tuning for Large Language Models
- AgentCaster: Reasoning-Guided Tornado Forecasting
- MC#: Mixture Compressor for Mixture-of-Experts Large Models
- A Layered Intuition -- Method Model with Scope Extension for LLM Reasoning
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- Failure-Driven Workflow Refinement
- Reinforcement Learning for Traversing Chemical Structure Space: Optimizing Transition States and Minimum Energy Paths of Molecules
- Improving AGI Evaluation: A Data Science Perspective
- Evaluating LLM-Based Process Explanations under Progressive Behavioral-Input Reduction
- Active Model Selection for Large Language Models
- What Is Your Agent's GPA? A Framework for Evaluating Agent Goal-Plan-Action Alignment
- GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
- From Description to Detection: LLM based Extendable O-RAN Compliant Blind DoS Detection in 5G and Beyond
- The fragility of "cultural tendencies" in LLMs
- FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
- CDTP: A Large-Scale Chinese Data-Text Pair Dataset for Comprehensive Evaluation of Chinese LLMs
- A Set of Quebec-French Corpus of Regional Expressions and Terms
- A Multidisciplinary Design and Optimization (MDO) Agent Driven by Large Language Models
- Benchmarking Open-Source Large Language Models for Persian in Zero-Shot and Few-Shot Learning
- COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- Toward a unified framework for data-efficient evaluation of large language models
- LLM Chemistry Estimation for Multi-LLM Recommendation
- PoseGaze-AHP: A Knowledge-Based 3D Dataset for AI-Driven Ocular and Postural Diagnosis
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
- EvoSpeak: Large Language Models for Interpretable Genetic Programming-Evolved Heuristics
- Graph-S3: Enhancing Agentic textual Graph Retrieval with Synthetic Stepwise Supervision
- Color2Struct: efficient and accurate deep-learning inverse design of structural color with controllable inference
- RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation
- QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs
- Benchmarking ECG Foundational Models: A Reality Check Across Clinical Tasks
- OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment
- Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
- SCI-Verifier: Scientific Verifier with Thinking
- Chat to Chip: Large Language Model Based Design of Arbitrarily Shaped Metasurfaces
- The problem of alignment
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in Large Language Model Routing
- Act as an expert in psychometry. The evaluation of large language models utility in psychological tests cross-cultural adaptations
- Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs
- Smoothing-Based Conformal Prediction for Balancing Efficiency and Interpretability
- A model of errors in transformers
- Mixture of Detectors: A Compact View of Machine-Generated Text Detection
- KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI
- An LLM-Powered Agent for Real-Time Analysis of the Vietnamese IT Job Market
- SAGE: A Realistic Benchmark for Semantic Understanding
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
- UniTransfer: Video Concept Transfer via Progressive Spatial and Timestep Decomposition
- Integrated Framework for LLM Evaluation with Answer Generation
- SSTAG: Structure-Aware Self-Supervised Learning Method for Text-Attributed Graphs
- DocFetch - Towards Generating Software Documentation from Multiple Software Artifacts
- <i>ToxTempAssistant</i> : using large language models to standardise cell-based toxicological test method descriptions
- Algebraic Approach to Ridge-Regularized Mean Squared Error Minimization in Minimal ReLU Neural Network
- Advancing Thesaurus Construction With Generative AI: A Structured Approach to Synonym Identification and Scope Note Development
- AECBench: A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field
- Agentic AutoSurvey: Let LLMs Survey LLMs
- Prompt-in-Content Attacks: Exploiting Uploaded Inputs to Hijack LLM Behavior
- Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models
- Bounded PCTL Model Checking of Large Language Model Outputs
- MEF: A Systematic Evaluation Framework for Text-to-Image Models
- A Knowledge Graph-based Retrieval-Augmented Generation Framework for Algorithm Selection in the Facility Layout Problem
- SilentStriker:Toward Stealthy Bit-Flip Attacks on Large Language Models
- SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
- How ChatGPT and Gemini View the Elements of Communication Competence of Large Language Models: A Pilot Study
- Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment
- Controlled Yet Natural: A Hybrid BDI-LLM Conversational Agent for Child Helpline Training
- Evaluating LLM Generated Detection Rules in Cybersecurity
- Improving Deep Tabular Learning
- Psychology's Questionable Research Fundamentals (QRFs): Key problems in quantitative psychology and psychological measurement beyond Questionable Research Practices (QRPs)
- Defining and Monitoring Complex Robot Activities via LLMs and Symbolic Reasoning
- Simulating a Bias Mitigation Scenario in Large Language Models
- DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models
- An Empirical Analysis of VLM-based OOD Detection: Mechanisms, Advantages, and Sensitivity
- When Large Language Models Meet UAV Projects: An Empirical Study from Developers' Perspective
- Multi Anatomy X-Ray Foundation Model
- Towards Deeper Understanding of Natural User Interactions in Virtual Reality Based Assembly Tasks
- Evaluating undergraduate mathematics examinations in the era of generative AI: a curriculum-level case study
- Large Language Models for Security Operations Centers: A Comprehensive Survey
- MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
- Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
- QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments
- Adapting Vision-Language Models for Neutrino Event Classification in High-Energy Physics
- Investigating Student Interaction Patterns with Large Language Model-Powered Course Assistants in Computer Science Courses
- No-Knowledge Alarms for Misaligned LLMs-as-Judges
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
- Stack Overflow Is Not Dead Yet: Crowd Answers Still Matter
- Foundational Models and Federated Learning: Survey, Taxonomy, Challenges and Practical Insights
- Artificial intelligence for representing and characterizing quantum systems
- AFD-SLU: Adaptive Feature Distillation for Spoken Language Understanding
- A Foundation Model for Chest X-ray Interpretation with Grounded Reasoning via Online Reinforcement Learning
- Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning
- LLM-GUARD: Large Language Model-Based Detection and Repair of Bugs and Security Vulnerabilities in C++ and Python
- Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
- Txt2Sce: Scenario Generation for Autonomous Driving System Testing Based on Textual Reports
- Federated Foundation Models in Harsh Wireless Environments: Prospects, Challenges, and Future Directions
- The Fools are Certain; the Wise are Doubtful: Exploring LLM Confidence in Code Completion
- GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
- ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts
- Fine-Tuning Vision-Language Models for Neutrino Event Analysis in High-Energy Physics Experiments
- MATRIX: Multi-Agent simulaTion fRamework for safe Interactions and conteXtual clinical conversational evaluation
- DELIVER: A System for LLM-Guided Coordinated Multi-Robot Pickup and Delivery using Voronoi-Based Relay Planning
- Beyond Benchmark: LLMs Evaluation with an Anthropomorphic and Value-oriented Roadmap
- PerFairX: Is There a Balance Between Fairness and Personality in Large Language Model Recommendations?
- The Promise of Large Language Models in Digital Health: Evidence from Sentiment Analysis in Online Health Communities
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
- Clean Code, Better Models: Enhancing LLM Performance with Smell-Cleaned Dataset
- WebGeoInfer: A Structure-Free and Multi-Stage Framework for Geolocation Inference of Devices Exposing Information
- Advancing Autonomous Incident Response: Leveraging LLMs and Cyber Threat Intelligence
- Exploring the Potential of Large Language Models in Fine-Grained Review Comment Classification
- ReqInOne: A Large Language Model-Based Agent for Software Requirements Specification Generation
- CS-Agent: LLM-based Community Search via Dual-agent Collaboration
- SinLlama -- A Large Language Model for Sinhala
- LPGNet: A Lightweight Network with Parallel Attention and Gated Fusion for Multimodal Emotion Recognition
- ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
- SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
- RAGTrace: Understanding and Refining Retrieval-Generation Dynamics in Retrieval-Augmented Generation
- Fast, Convex and Conditioned Network for Multi-Fidelity Vectors and Stiff Univariate Differential Equations
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
- Hierarchical Text Classification Using Black Box Large Language Models
- CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
- Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
- Win-k: Improved Membership Inference Attacks on Small Language Models
- DAMR: Efficient and Adaptive Context-Aware Knowledge Graph Question Answering with LLM-Guided MCTS
- Objective Metrics for Evaluating Large Language Models Using External Data Sources
- Comparison of Large Language Models for Deployment Requirements
- On LLM-Assisted Generation of Smart Contracts from Business Processes
Related