Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
2020/04/15 by Miles Brundage, Shahar Avin, Brundage, Miles +123 · 2 voices · 42 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Ethics and Social Impacts of AI #Law, AI, and Intellectual Property #cs.CY
paper · pdf · doi:10.48550/arxiv.2004.07213
openalex publication_date 2020/04/15 · arxiv created 2020/04/20 · arxiv updated 2020/04/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
With the recent wave of progress in artificial intelligence (AI) has come a growing awareness of the large-scale impacts of AI systems, and recognition that existing regulations and norms in industry and academia are insufficient to ensure responsible AI development. In order for AI developers to earn trust from system users, customers, civil society, governments, and other stakeholders that they are building AI responsibly, they will need to make verifiable claims to which they can be held accountable. Those outside of a given organization also need effective means of scrutinizing such claims. This report suggests various steps that different stakeholders can take to improve the verifiability of claims made about AI systems and their associated development processes, with a focus on providing evidence about the safety, security, fairness, and privacy protection of AI systems. We analyze ten mechanisms for this purpose--spanning institutions, software, and hardware--and make recommendations aimed at implementing, exploring, or improving those mechanisms.
Cited by
- Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management
- Assessing High-Risk AI Systems under the EU AI Act: From Legal Requirements to Technical Verification
- The 2025 Foundation Model Transparency Index
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models
- From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
- "Show Me You Comply... Without Showing Me Anything": Zero-Knowledge Software Auditing for AI-Enabled Systems
- Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
- Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles
- From misinformation to climate crisis: Navigating vulnerabilities in the cyber-physical-social systems
- Measuring What Matters: Connecting AI Ethics Evaluations to System Attributes, Hazards, and Harms
- Emergent evaluation hubs in a decentralizing large language model ecosystem
- The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies
- AI Ethics Education in India: A Syllabus-Level Review of Computing Courses
- What Makes LLM Agent Simulations Useful for Policy Practice? An Iterative Design Study in Emergency Preparedness
- Explainable Graph Neural Networks: Understanding Brain Connectivity and Biomarkers in Dementia
- The Aegis Protocol: A Foundational Security Framework for Autonomous AI Agents
- Enhancing XAI Interpretation through a Reverse Mapping from Insights to Visualizations
- TechOps: Technical Documentation Templates for the AI Act
- Towards Transparent Ethical AI: A Roadmap for Trustworthy Robotic Systems
- AutoSIGHT: Automatic Eye Tracking-based System for Immediate Grading of Human experTise
- Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques
- The Hidden Costs of AI: A Review of Energy, E-Waste, and Inequality in Model Development
- Strategic Alignment Patterns in National AI Policies
- Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
- AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
- Black-Box Access is Insufficient for Rigorous AI Audits
- Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs
- Bridging the Artificial Intelligence Governance Gap: The United States' and China's Divergent Approaches to Governing General-Purpose Artificial Intelligence
- HADA: Human-AI Agent Decision Alignment Architecture
- Machine vs Machine: Using AI to Tackle Generative AI Threats in Assessment
- Bottom-Up Perspectives on AI Governance: Insights from User Reviews of AI Products
- Risks of AI-driven product development and strategies for their mitigation
- What Is AI Safety? What Do We Want It to Be?
- Third-party compliance reviews for frontier AI safety frameworks
- Toward a Science of Intent: Closure Gaps and Delegation Envelopes for Open-World AI Agents
- Coupled Control, Structured Memory, and Verifiable Action in Agentic AI (SCRAT -- Stochastic Control with Retrieval and Auditable Trajectories): A Comparative Perspective from Squirrel Locomotion and Scatter-Hoarding
- Filling gaps in trustworthy development of AI
- Enhancing Trust Through Standards: A Comparative Risk-Impact Framework for Aligning ISO AI Standards with Global Ethical and Regulatory Contexts
- Contemplative Agent
- Framework, Standards, Applications and Best practices of Responsible AI : A Comprehensive Survey
- Towards a robust and trustworthy machine learning system development: An engineering perspective
Discussions
Related