vix.ing · top · new · best · stats · spec

Minimum Levels of Interpretability for Artificial Moral Agents

2023/07/02 by Avish Vijayaraghavan, Vijayaraghavan, Avish, Cosmin Badea +1 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences

paper · pdf · doi:10.48550/arxiv.2307.00660

openalex publication_date 2023/07/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01

Abstract

As artificial intelligence (AI) models continue to scale up, they are becoming more capable and integrated into various forms of decision-making systems. For models involved in moral decision-making, also known as artificial moral agents (AMA), interpretability provides a way to trust and understand the agent's internal reasoning mechanisms for effective use and error correction. In this paper, we provide an overview of this rapidly-evolving sub-field of AI interpretability, introduce the concept of the Minimum Level of Interpretability (MLI) and recommend an MLI for various types of agents, to aid their safe deployment in real-world settings.

Cited by

Related