vix.ing · top · new · best · stats · spec

OpenEthics: A Comprehensive Ethical Evaluation of Open-Source Generative Large Language Models

2025/05/21 by Yıldırım Özen, Çetin, Burak Erinç, Burak Erinç Çetin +7
Computer Science · Medicine · Social Sciences · #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #Ethics and Social Impacts of AI #FOS: Computer and information sciences #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2505.16036

openalex publication_date 2025/05/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04

Abstract

Generative large language models present significant potential but also raise critical ethical concerns, including issues of safety, fairness, robustness, and reliability. Most existing ethical studies, however, are limited by their narrow focus, a lack of language diversity, and an evaluation of a restricted set of models. To address these gaps, we present a broad ethical evaluation of 29 recent open-source LLMs using a novel dataset that assesses four key ethical dimensions: robustness, reliability, safety, and fairness. Our analysis includes both a high-resource language, English, and a low-resource language, Turkish, providing a comprehensive assessment and a guide for safer model development. Using an LLM-as-a-Judge methodology, our experimental results indicate that many open-source models demonstrate strong performance in safety, fairness, and robustness, while reliability remains a key concern. Ethical evaluation shows cross-linguistic consistency, and larger models generally exhibit better ethical performance. We also show that jailbreak templates are ineffective for most of the open-source models examined in this study. We share all materials including data and scripts at https://github.com/metunlp/openethics

Citations

Related