LLM Ethics Benchmark: A Three-Dimensional Assessment System for Evaluating Moral Reasoning in Large Language Models
2025/05/01 by Junfeng Jiao, Saleh Afroogh, Jiao, Junfeng +8 · 11 citations
Social Sciences · #Artificial Intelligence in Law #Computers and Society (cs.CY) #FOS: Computer and information sciences
paper · pdf · doi:10.48550/arxiv.2505.00853
openalex publication_date 2025/05/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
This study establishes a novel framework for systematically evaluating the moral reasoning capabilities of large language models (LLMs) as they increasingly integrate into critical societal domains. Current assessment methodologies lack the precision needed to evaluate nuanced ethical decision-making in AI systems, creating significant accountability gaps. Our framework addresses this challenge by quantifying alignment with human ethical standards through three dimensions: foundational moral principles, reasoning robustness, and value consistency across diverse scenarios. This approach enables precise identification of ethical strengths and weaknesses in LLMs, facilitating targeted improvements and stronger alignment with societal values. To promote transparency and collaborative advancement in ethical AI development, we are publicly releasing both our benchmark datasets and evaluation codebase at https://github.com/ The-Responsible-AI-Initiative/LLMEthicsBenchmark.git.
Citations
- FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation
- SafetyBench: Evaluating the Safety of Large Language Models
- CMB: A Comprehensive Medical Benchmark in Chinese
- Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- MMBench: Is Your Multi-modal Model an All-around Player?
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- TrustGPT: A Benchmark for Trustworthy and Responsible Large Language Models
- LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models
- LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
- Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge Evaluation
- M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models
- Cue-CoT: Chain-of-thought Prompting for Responding to In-depth Dialogue Questions with LLMs
- C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
- AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
- GPT-4 Technical Report
- Large Language Models Encode Clinical Knowledge
- Large language models encode clinical knowledge
- Language Models as Agent Models
- GLUE-X: Evaluating Natural Language Understanding Models from an Out-of-distribution Generalization Perspective
- Probing Pre-Trained Language Models for Cross-Cultural Differences in Values
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- On the Opportunities and Risks of Foundation Models
- Systematic Evaluation of Causal Discovery in Visual Model Based Reinforcement Learning
- Measuring Coding Challenge Competence With APPS
- Dynabench: Rethinking Benchmarking in NLP
- CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review
- Measuring Mathematical Problem Solving With the MATH Dataset
- On the Dangers of Stochastic Parrots
- Calibrate Before Use: Improving Few-Shot Performance of Language Models
- Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences
- Logic-guided Semantic Representation Learning for Zero-Shot Relation Classification
- Language Models are Few-Shot Learners
- Artificial Intelligence, Values and Alignment
- Closing the AI Accountability Gap: Defining an End-to-End Framework for Internal Algorithmic Auditing
- Closing the AI accountability gap
- Social Bias Frames: Reasoning about Social and Power Implications of Language
- Artificial Intelligence: the global landscape of ethics guidelines
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Neural Legal Judgment Prediction in English
- Multi-Task Deep Neural Networks for Natural Language Understanding
- The Moral Machine experiment
- Gender Bias in Coreference Resolution
- Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems
- Concrete Problems in AI Safety
- The emotional dog and its rational tail: A social intuitionist approach to moral judgment.
- PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Cited by
Related