Model Cards for Model Reporting
2018/10/05 by Margaret Mitchell, Simone Wu, Andrew Zaldivar +6 · 6 voices · 198 citations
Social Sciences · Computer Science · #Ethics and Social Impacts of AI #Artificial Intelligence in Law #Privacy-Preserving Technologies in Data
paper · pdf · doi:10.1145/3287560.3287596
Abstract
Trained machine learning models are increasingly used to perform high-impact tasks in areas such as law enforcement, medicine, education, and employment. In order to clarify the intended use cases of machine learning models and minimize their usage in contexts for which they are not well suited, we recommend that released models be accompanied by documentation detailing their performance characteristics. In this paper, we propose a framework that we call model cards, to encourage such transparent model reporting. Model cards are short documents accompanying trained machine learning models that provide benchmarked evaluation in a variety of conditions, such as across different cultural, demographic, or phenotypic groups (e.g., race, geographic location, sex, Fitzpatrick skin type [15]) and intersectional groups (e.g., age and race, or sex and Fitzpatrick skin type) that are relevant to the intended application domains. Model cards also disclose the context in which models are intended to be used, details of the performance evaluation procedures, and other relevant information. While we focus primarily on human-centered machine learning models in the application fields of computer vision and natural language processing, this framework can be used to document any trained machine learning model. To solidify the concept, we provide cards for two supervised models: One trained to detect smiling faces in images, and one trained to detect toxic comments in text. We propose model cards as a step towards the responsible democratization of machine learning and related artificial intelligence technology, increasing transparency into how well artificial intelligence technology works. We hope this work encourages those releasing trained machine learning models to accompany model releases with similar detailed evaluation numbers and other relevant documentation.
Citations
Cited by
- Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
- Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification
- When Model Release Meets Model Reuse: Producer-Consumer Misalignment in Hugging Face
- What AI Red-Team Evaluations Can and Cannot Prove
- Unsafe at any AUC: Unlearned Lessons From Sociotechnical Disasters for Responsible AI
- STeMP: Spatio-Temporal Modelling Protocol
- Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator
- An Auditable Policy-Simulation Framework for Student Dropout in Intervention-Free Data
- Integrating High-Level Requirements to Low-Level Tests with Machine-Readable V&V Specifications
- The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits
- No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason
- Participatory provenance as representational auditing for AI-mediated public consultation
- A Large-Scale Measurement of AI Bill of Materials Completeness in Hugging Face Models
- Beyond Visibility and Technical Reuse: Public Application Transformation in Open-Source Model Ecosystems
- Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI
- A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance
- When Not to Automate: A Formal Protocol for Human Preservation in AI-Optimized Organizations
- Align AI to Dynamic Human-AI Workflows
- FALCON-Discover: Discovering Concentrated False-Confidence Regions for Calibration
- Human-AI governance (HAIG): A trust-utility approach
- Building an Open AIBOM Standard in the Wild: An Experience Report on Extending the SPDX SBOM (ISO/IEC 5962:2021) for AI Supply Chains
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- Competing visions of ethical AI: a case study of OpenAI
- A Human-Centric Framework for Data Attribution in Large Language Models
- People Can Accurately Predict Behavior of Complex Algorithms That Are Available, Compact, and Aligned
- Critical, but constructive: defining, detecting, and addressing bias in Computational Social Science
- User Privacy and Large Language Models: An Analysis of Frontier Developers' Privacy Policies
- Bonsai: Intentional and Personalized Social Media Feeds
- Anwendungen, Herausforderungen und ein vertrauenswürdiger Umgang mit künstlicher Intelligenz im Bereich Public Health
- Misinformation by Omission: The Need for More Environmental Transparency in AI
- An Interdisciplinary Approach to Human-Centered Machine Translation
- Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
- Provocations from the Humanities for Generative AI Research
- Misclassification in Automated Content Analysis Causes Bias in Regression. Can We Fix It? Yes We Can!
- Measuring complex constructs in large-scale text with computational social mixed methods
- Educational Strategies for Clinical Supervision of Artificial Intelligence Use
- Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
- Between fact and fairy: tracing the hallucination metaphor in AI discourse
- What Matters in Deep Learning for Time Series Forecasting?
- Essential work, invisible workers: The role of digital curation in <scp>COVID</scp>‐19 Open Science
- Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management
- Adaptive Accountability in Networked MAS: Tracing and Mitigating Emergent Norms at Scale
- Toward Ethical AI Through Bayesian Uncertainty in Neural Question Answering
- Best Practices For Empirical Meta-Algorithmic Research: Guidelines from the COSEAL Research Network
- Smart Data Portfolios: A Governance Framework for AI Training Data
- Prompt Governance? On Governing Technologies Governed by Natural Language
- Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Stud
- Towards Open Standards for Systemic Complexity in Digital Forensics
- SafeGen: Embedding Ethical Safeguards in Text-to-Image Generation
- AI Transparency Atlas: Framework, Scoring, and Real-Time Model Card Evaluation Pipeline
- The 2025 Foundation Model Transparency Index
- Auto-BenchmarkCard: Automated Synthesis of Benchmark Documentation
- An Empirical Framework for Evaluating Semantic Preservation Using Hugging Face
- A Unifying Human-Centered AI Fairness Framework
- "Having Confidence in My Confidence Intervals": How Data Users Engage with Privacy-Protected Wikipedia Data
- Industrial AI Robustness Card: Evaluating and Monitoring Time Series Models
- Eval Factsheets: A Structured Framework for Documenting AI Evaluations
- Institutional AI Sovereignty Through Gateway Architecture: Implementation Report from Fontys ICT
- Enhancing Transparency and Traceability in Healthcare AI: The AI Product Passport
- Whose Personae? Synthetic Persona Experiments in LLM Research and Pathways to Transparency
- AI/ML Model Cards in Edge AI Cyberinfrastructure: towards Agentic AI
- Mirror, Mirror on the Wall -- Which is the Best Model of Them All?
- Domain-constrained Synthesis of Inconsistent Key Aspects in Textual Vulnerability Descriptions
- Identifying the Supply Chain of AI for Trustworthiness and Risk Management in Critical Applications
- From Machine Learning Documentation to Requirements: Bridging Processes with Requirements Languages
- MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm
- Bias and Fairness in Large Language Models: A Survey
- Competing narratives in AI ethics: a defense of sociotechnical pragmatism
- AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
- Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
- Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
- EvalCards: A Framework for Standardized Evaluation Reporting
- AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda
- AgentSLA : Towards a Service Level Agreement for AI Agents
- Community Detection on Model Explanation Graphs for Explainable AI
- Enhancing software product lines with machine learning components
- From the Rock Floor to the Cloud: A Systematic Survey of State-of-the-Art NLP in Battery Life Cycle
- Can AI be Accountable?
- F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill
- Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
- Six Human-Centered Artificial Intelligence Grand Challenges
- Position: Evaluation Scores Are Perishable Knowledge Claims
- When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
- ROADMAP: An Ontology of Medical AI Models and Datasets
- Metrics for Artificial Intelligence in Medicine: A Reference Resource
- Exploring the Intersection of AI, Language, and Law: A Bibliometric Analysis
- Metaethical perspectives on ‘benchmarking’ AI ethics
- Equipping Speech-Language Clinicians for the Critical Appraisal of an Artificial Intelligence–Driven, Evidence-Based Future
- An Adaptive Responsible AI Governance Framework for Decentralized Organizations
- Power to the People? Opportunities and Challenges for Participatory AI
- Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
- The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
- Teaching Machine Learning to Software Engineers
- AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
- Metacognition Should Be the Scientific Framework for Bounded and Effective Self-Governance in Generative AI
- Charting the Sociotechnical Gap in Explainable AI: A Framework to Address the Gap in XAI
- Trustworthy AI: From Principles to Practices
- The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
- Segment Anything
- Large language models encode clinical knowledge
- Integrating Machine Learning Standards in Disseminating Machine Learning Research
- The use of LLMs to annotate data in management research: Foundational guidelines and warnings
- Comparison of generalised additive models and neural networks in applications: A systematic review
- Policy Cards: Machine-Readable Runtime Governance for Autonomous AI Agents
- Ethics in Linguistics
- Mutual Wanting in Human--AI Interaction: Empirical Evidence from Large-Scale Analysis of GPT Model Transitions
- Reasoning About Reasoning: Towards Informed and Reflective Use of LLM Reasoning in HCI
- LLM-augmented empirical game theoretic simulation for social-ecological systems
- HIKMA: Human-Inspired Knowledge by Machine Agents through a Multi-Agent Framework for Semi-Autonomous Scientific Conferences
- Embedding Explainable AI in NHS Clinical Safety: The Explainability-Enabled Clinical Safety Framework (ECSF)
- Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
- What do model reports say about their ChemBio benchmark evaluations? Comparing recent releases to the STREAM framework
- Machine Learning Practices Outside Big Tech: How Resource Constraints Challenge Responsible Development
- Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
- Speculative Model Risk in Healthcare AI: Using Storytelling to Surface Unintended Harms
- Machine Learning and Public Health: Identifying and Mitigating Algorithmic Bias through a Systematic Review
- JEDA: Query-Free Clinical Order Search from Ambient Dialogues
- On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy
- SeFEF: A Seizure Forecasting Evaluation Framework
- Reverse Supervision at Scale: Exponential Search Meets the Economics of Annotation
- SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
- Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation
- Enabling Responsible, Secure and Sustainable Healthcare AI - A Strategic Framework for Clinical and Operational Impact
- Towards Meaningful Transparency in Civic AI Systems
- Measuring What Matters: The AI Pluralism Index
- Human-aligned AI Model Cards with Weighted Hierarchy Architecture
- TAIBOM: Bringing Trustworthiness to AI-Enabled Systems
- Assessing Human Rights Risks in AI: A Framework for Model Evaluation
- Accountability Capture: How Record-Keeping to Support AI Transparency and Accountability (Re)shapes Algorithmic Oversight
- A global log for medical AI
- Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification
- Machine Learning Workflows in Climate Modeling: Design Patterns and Insights from Case Studies
- Emergent evaluation hubs in a decentralizing large language model ecosystem
- Agentic Services Computing
- Exploring Opportunities to Support Novice Visual Artists' Inspiration and Ideation with Generative AI
- ML-Asset Management: Curation, Discovery, and Utilization
- Transparency of medical artificial intelligence systems
- Culture machine: How MetaCLIP codifies culture
- "I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work
- What Is The Political Content in LLMs' Pre- and Post-Training Data?
- Does AI Coaching Prepare us for Workplace Negotiations?
- Longitudinal Monitoring of LLM Content Moderation of Social Issues
- Measurements, Algorithms, and Presentations of Reality: Framing Interactions with AI-Enabled Decision Support
- Scaling, Lock-In, and Proxy Compliance: A Political Economy of Responsible AI
- ContextNest: Verifiable Context Governance for Autonomous AI Agent
- Investigating machine moral judgement through the Delphi experiment
- Advancing neurotech justice in youth digital mental health: insights from an interdisciplinary and cross-generational workshop
- Bureaucratic Silences: What the Canadian AI Register Reveals, Omits, and Obscures
- Inspectable AI for Science: A Research Object Approach to Generative AI Governance
- A Comprehensive Review of Real-Time Multi-View Multi-Person Markerless Motion Capture
- Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and Considerations
- Open-Source Large Language Models in Education: A Narrative Review of Evidence, Pedagogical Roles, and Learning Outcomes
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- A Feminist Account of Intersectional Algorithmic Fairness
- Auditing large language models: a three-layered approach
- Archival Paradata and Artificial Intelligence in Archaeology
- Pensar con la mirada: cocreación humano-algorítmica y agencia distribuida en Visions of Destruction
- From moral panic to pragmatic governance: reframing AI’s societal impacts in employment, education, and ethics
- Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling
- "It Was a Magical Box": Understanding Practitioner Workflows and Needs in Optimization
- QUINTA: Reflexive Sensibility For Responsible AI Research and Data-Driven Processes
- Right-to-Override for Critical Urban Control Systems: A Deliberative Audit Method for Buildings, Power, and Transport
- Practitioners' Perspectives on a Differential Privacy Deployment Registry
- Prompt Commons: Collective Prompting as Governance for Urban AI
- Collective Recourse for Generative Urban Visualizations
- Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- Closing the AI accountability gap
- mFARM: Towards Multi-Faceted Fairness Assessment based on HARMs in Clinical Decision Support
- Gaming and Cooperation in Federated Learning: What Can Happen and How to Monitor It
- Computational Social Science and Critical Studies of Education and Technology: An Improbable Combination?
- Machine Learning for Medicine Must Be Interpretable, Shareable, Reproducible and Accountable by Design
- Who Owns The Robot?: Four Ethical and Socio-technical Questions about Wellbeing Robots in the Real World through Community Engagement
- One VLM, Two Roles: Stage-Wise Routing and Specialty-Level Deployment for Clinical Workflows
- The AI Model Risk Catalog: What Developers and Researchers Miss About Real-World AI Harms
- A.I. Robustness: a Human-Centered Perspective on Technological Challenges and Opportunities
- Trust and trustworthy artificial intelligence: A research agenda for AI in the environmental sciences
- "She was useful, but a bit too optimistic": Augmenting Design with Interactive Virtual Personas
- Toward Operationalizing Pipeline-aware ML Fairness: A Research Agenda for Developing Practical Guidelines and Tools
- Impact Assessment Card: Communicating Risks and Benefits of AI Uses
- Breakable Machine: A K-12 Classroom Game for Transformative AI Literacy Through Spoofing and eXplainable AI (XAI)
- The AI-Fraud Diamond: A Novel Lens for Auditing Algorithmic Deception
- Documenting Deployment with Fabric: A Repository of Real-World AI Governance
- PTMPicker: Facilitating Efficient Pretrained Model Selection for Application Developers
- Bias is a Math Problem, AI Bias is a Technical Problem: 10-year Literature Review of AI/LLM Bias Research Reveals Narrow [Gender-Centric] Conceptions of 'Bias', and Academia-Industry Gap
- STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
- Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
- TechOps: Technical Documentation Templates for the AI Act
- Towards Experience-Centered AI: A Framework for Integrating Lived Experience in Design and Development
- Anatomy of a Machine Learning Ecosystem: 2 Million Models on Hugging Face
- EICAP: Deep Dive in Assessment and Enhancement of Large Language Models in Emotional Intelligence through Multi-Turn Conversations
- Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
- I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
- Co-Producing AI: Toward an Augmented, Participatory Lifecycle
- A Survey on Deep Multi-Task Learning in Connected Autonomous Vehicles
- Objectifying the Subjective: Cognitive Biases in Topic Interpretations
- Towards Sustainability Model Cards
- OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
Discussions
- it's this paper, you'll notice @mmitchell.bsky.social and timnit who has me blocked so i can't @ her as the most famous names on it arxiv.org/abs/1810.03993 [bsky, 23 points, 1 comments]
- For those unfamiliar, they’re a tool from @mmitchell.bsky.social et al. for documenting ML models (and their social considerations). https://arxiv.org/abs/1810.03993 [bsky, 7 points, 0 comments]
- If (like me) you wondered what the origin of system cards (aka model cards) was…
arxiv.org/pdf/1810.03993
(Kinda wish all software came with a doc like this!) [bsky, 5 points, 1 comments]
- As it so happens, I am the world's leading expert on model cards (!) Actually I "invented"* them ca. 2018. arxiv.org/abs/1810.03993 I have also operationalised these at Google ("closed" company) and H [bsky, 1 points, 1 comments]
- The paper "Model Cards for Model Reporting" presents a framework for documenting ML models, stressing the need for cards detailing performance across diverse demographics to foster responsible AI use. [bsky, 0 points, 0 comments]
- They will learn about the possibilities & limits of fairness metrics using some known datasets. It prepares them for documenting their own model with a model card as formulated by @mmitchell.bsky.soci [bsky, 0 points, 1 comments]
Related