Datasheets for datasets
2021/11/19 by Timnit Gebru, Jamie Morgenstern, Briana Vecchione +5 · 195 citations
Decision Sciences · Computer Science · #Data Quality and Management #Data Stream Mining Techniques #Machine Learning and Data Classification
paper · doi:10.1145/3458723
Abstract
Documentation to facilitate communication between dataset creators and consumers.
Cited by
- Associations Between Support-Seekers' Cross-Community Interactions and Their Engagement with Received Comments in Online Health Communities
- Fairness and Abstraction in Sociotechnical Systems
- AuditRepairBench: A Paired-Execution Trace Corpus for Evaluator-Channel Ranking Instability in Agent Repair
- When Model Release Meets Model Reuse: Producer-Consumer Misalignment in Hugging Face
- Unsafe at any AUC: Unlearned Lessons From Sociotechnical Disasters for Responsible AI
- Spaghetti Architect: A Contamination-Resistant, By-Construction-Labelled, Multi-Language Code Dataset Generator
- The Aura in the Machine: Genealogy and the Status of the Work of Art in the Generative Era
- Integrating High-Level Requirements to Low-Level Tests with Machine-Readable V&V Specifications
- The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits
- Participatory provenance as representational auditing for AI-mediated public consultation
- Beyond Visibility and Technical Reuse: Public Application Transformation in Open-Source Model Ecosystems
- Posts of Peril: Detecting Information About Hazards in Text
- A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance
- NAMESAKES: Probing Identity Memorization in Text-to-Image Models
- TikStance: A Multimodal and Hierarchical Dataset for Multi-target Stance Analysis in TikTok Political Conversations
- Align AI to Dynamic Human-AI Workflows
- World Wide Models: Literary Tools for Cultural AI
- FALCON-Discover: Discovering Concentrated False-Confidence Regions for Calibration
- Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR
- Building an Open AIBOM Standard in the Wild: An Experience Report on Extending the SPDX SBOM (ISO/IEC 5962:2021) for AI Supply Chains
- Competing visions of ethical AI: a case study of OpenAI
- A Human-Centric Framework for Data Attribution in Large Language Models
- What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification
- Anwendungen, Herausforderungen und ein vertrauenswürdiger Umgang mit künstlicher Intelligenz im Bereich Public Health
- Potemkin Understanding in Large Language Models
- A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
- Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
- Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking
- Provocations from the Humanities for Generative AI Research
- The Cake that is Intelligence and Who Gets to Bake it: An AI Analogy and its Implications for Participation
- Multi-Platform Aggregated Dataset of Online Communities (MADOC)
- GoogleTrendArchive: A Year-Long Archive of Real-Time Web Search Trends Worldwide
- Trust and Transparency in Contact Tracing Applications
- Identifying Bias in AI using Simulation
- Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI
- Measuring complex constructs in large-scale text with computational social mixed methods
- MUTE: Data-Similarity Driven Multi-hot Target Encoding for Neural Network Design
- About Face: A Survey of Facial Recognition Evaluation
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
- Nose to Glass: Looking In to Get Beyond
- Principles to Practices for Responsible AI: Closing the Gap
- Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus
- Mitigating Dataset Harms Requires Stewardship: Lessons from 1000 Papers
- Show Your Work: Improved Reporting of Experimental Results
- Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure
- Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It
- Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning
- Between Subjectivity and Imposition: Power Dynamics in Data Annotation for Computer Vision
- Do Current Retrievers Cover All the Evidence? A Controlled Study of Conjunctive Cross-Page Retrieval
- The Gray Area: Characterizing Moderator Disagreement on Reddit
- On the Opportunities and Risks of Foundation Models
- BiToD: A Bilingual Multi-Domain Dataset For Task-Oriented Dialogue Modeling
- Medical Imaging AI Competitions Lack Fairness
- XOR QA: Cross-lingual Open-Retrieval Question Answering
- Network Analysis of Cyberbullying Interactions on Instagram
- Smart Data Portfolios: A Governance Framework for AI Training Data
- ModelTables: A Corpus of Tables about Models
- SNIC: Synthesized Noisy Images using Calibration
- Evaluation of AI Ethics Tools in Language Models: A Developers' Perspective Case Stud
- Open Source Software and Data for Human Service Development: A Case Study on Predicting Housing Instability
- SafeGen: Embedding Ethical Safeguards in Text-to-Image Generation
- Textual Data Bias Detection and Mitigation -- An Extensible Pipeline with Experimental Evaluation
- AI Didn't Start the Fire: Examining the Stack Exchange Moderator and Contributor Strike
- FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
- A Unifying Human-Centered AI Fairness Framework
- "Having Confidence in My Confidence Intervals": How Data Users Engage with Privacy-Protected Wikipedia Data
- LLM Harms: A Taxonomy and Discussion
- Eval Factsheets: A Structured Framework for Documenting AI Evaluations
- The State's Politics of "Fake Data"
- GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization
- Forced Migration and Information-Seeking Behavior on Wikipedia: Insights from the Ukrainian Refugee Crisis
- Whose Personae? Synthetic Persona Experiments in LLM Research and Pathways to Transparency
- Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
- From Machine Learning Documentation to Requirements: Bridging Processes with Requirements Languages
- Bias in, Bias out: Annotation Bias in Multilingual Large Language Models
- Understanding the Complexities of Responsibly Sharing NSFW Content Online
- Patterns Count-Based Labels for Datasets
- AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework
- Non-portability of Algorithmic Fairness in India
- mmJEE-Eval: A Bilingual Multimodal Benchmark for Evaluating Scientific Reasoning in Vision-Language Models
- Semantic maps and metrics for science Semantic maps and metrics for science using deep transformer encoders
- QuAnTS: Question Answering on Time Series
- Who Evaluates AI's Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations
- What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
- miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward
- EvalCards: A Framework for Standardized Evaluation Reporting
- AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Before the Clinic: Transparent and Operable Design Principles for Healthcare AI
- F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill
- WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing
- Uncertainty as a Form of Transparency: Measuring, Communicating, and Using Uncertainty
- Representation Matters: Assessing the Importance of Subgroup Allocations\n in Training Data
- The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics
- Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
- Position: Evaluation Scores Are Perishable Knowledge Claims
- When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- MERLOT: Multimodal Neural Script Knowledge Models
- EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
- It's COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks
- Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
- KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report
- The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
- Exploring Data Pipelines through the Process Lens: a Reference Model forComputer Vision
- Infini-News: Efficiently Queryable Access to 1.3 Billion Processed Common Crawl News Articles
- Human-Centered Explainable AI (XAI): From Algorithms to User Experiences
- Metacognition Should Be the Scientific Framework for Bounded and Effective Self-Governance in Generative AI
- PepSpecBench: A Unified Evaluation Benchmark for Peptide Tandem Mass Spectrometry Prediction
- Reproducibility in Machine Learning for Health
- SegmentMeIfYouCan: A Benchmark for Anomaly Segmentation
- Exploring the Intersection of AI, Language, and Law: A Bibliometric Analysis
- What Would Jiminy Cricket Do? Towards Agents That Behave Morally
- The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems
- Retrieval Collapses When AI Pollutes the Web
- MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
- If the archive can’t consent: Reimagining motion data and AI ethics for dance’s embodied histories
- Adaptive Data Collection for Latin-American Community-sourced Evaluation of Stereotypes (LACES)
- Sign Language Recognition, Generation, and Translation: An Interdisciplinary Perspective
- Reasoning About Reasoning: Towards Informed and Reflective Use of LLM Reasoning in HCI
- Stop the Nonconsensual Use of Nude Images in Research
- A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection
- HIKMA: Human-Inspired Knowledge by Machine Agents through a Multi-Agent Framework for Semi-Autonomous Scientific Conferences
- What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus
- Fair Generative Modeling via Weak Supervision
- VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
- BO4Mob: Bayesian Optimization Benchmarks for High-Dimensional Urban Mobility Problem
- DroneAudioset: An Audio Dataset for Drone-based Search and Rescue
- Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
- AFRICAPTION: Establishing a New Paradigm for Image Captioning in African Languages
- Predicting the Unpredictable: Reproducible BiLSTM Forecasting of Incident Counts in the Global Terrorism Database (GTD)
- Machine Learning and Public Health: Identifying and Mitigating Algorithmic Bias through a Systematic Review
- JEDA: Query-Free Clinical Order Search from Ambient Dialogues
- The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
- Towards Traceability in Data Ecosystems using a Bill of Materials Model
- Framing Unionization on Facebook: Communication around Representation Elections in the United States
- Enabling Responsible, Secure and Sustainable Healthcare AI - A Strategic Framework for Clinical and Operational Impact
- Measuring What Matters: The AI Pluralism Index
- Lean Finder: Semantic Search for Mathlib That Understands User Intents
- Human-aligned AI Model Cards with Weighted Hierarchy Architecture
- Assessing Human Rights Risks in AI: A Framework for Model Evaluation
- COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
- Accountability Capture: How Record-Keeping to Support AI Transparency and Accountability (Re)shapes Algorithmic Oversight
- Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification
- Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generation
- RoBiologyDataChoiceQA: A Romanian Dataset for improving Biology understanding of Large Language Models
- On Explaining Proxy Discrimination and Unfairness in Individual Decisions Made by AI Systems
- Agentic Services Computing
- Exploring Opportunities to Support Novice Visual Artists' Inspiration and Ideation with Generative AI
- Artificial Intelligence in Multimodal Learning Process Analytics
- VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
- Fostering Robots: A Governance-First Conceptual Framework for Domestic, Curriculum-Based Trajectory Collection
- MASH: A Multiplatform and Multimodal Annotated Dataset for Societal Impact of Hurricane
- Easy, Reproducible and Quality-Controlled Data Collection with Crowdaq
- "I Don't Think RAI Applies to My Model'' -- Engaging Non-champions with Sticky Stories for Responsible AI Work
- Does AI Coaching Prepare us for Workplace Negotiations?
- GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow
- What Does Your Benchmark Really Measure? A Framework for Robust Inference of AI Capabilities
- DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series
- ContextNest: Verifiable Context Governance for Autonomous AI Agent
- Bureaucratic Silences: What the Canadian AI Register Reveals, Omits, and Obscures
- Slicing the past to predict the future: Recasting data slicing as curatorial work in ML development and evaluation
- Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and Considerations
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- A Feminist Account of Intersectional Algorithmic Fairness
- Applying machine learning and AI for nanofiltration membranes water applications: a review
- From moral panic to pragmatic governance: reframing AI’s societal impacts in employment, education, and ethics
- WolBanking77: Wolof Banking Speech Intent Classification Dataset
- PoliTok-DE: A Multimodal Dataset of Political TikToks and Deletions From Germany
- "It Was a Magical Box": Understanding Practitioner Workflows and Needs in Optimization
- Modeling Worlds in Text
- QUINTA: Reflexive Sensibility For Responsible AI Research and Data-Driven Processes
- Assessing Historical Structural Oppression Worldwide via Rule-Guided Prompting of Large Language Models
- HumBugDB: A Large-scale Acoustic Mosquito Dataset
- Op-Fed: Opinion, Stance, and Monetary Policy Annotations on FOMC Transcripts Using Active Learning
- Right-to-Override for Critical Urban Control Systems: A Deliberative Audit Method for Buildings, Power, and Transport
- Practitioners' Perspectives on a Differential Privacy Deployment Registry
- Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
- Aligning AI With Shared Human Values
- Prompt Commons: Collective Prompting as Governance for Urban AI
- Question-Driven Design Process for Explainable AI User Experiences
- Collective Recourse for Generative Urban Visualizations
- Retiring Adult: New Datasets for Fair Machine Learning
- Standards in the Preparation of Biomedical Research Metadata: A Bridge2AI Perspective
- MetaRAG: Metamorphic Testing for Hallucination Detection in RAG Systems
- AI4D -- African Language Program
- Transparency of medical artificial intelligence systems
- Are LLMs Enough for Hyperpartisan, Fake, Polarized and Harmful Content Detection? Evaluating In-Context Learning vs. Fine-Tuning
- Overcoming Failures of Imagination in AI Infused System Development and Deployment
- Closing the AI accountability gap
- Gaming and Cooperation in Federated Learning: What Can Happen and How to Monitor It
- Computational Social Science and Critical Studies of Education and Technology: An Improbable Combination?
- MatPROV: A Provenance Graph Dataset of Material Synthesis Extracted from Scientific Literature
- Who Owns The Robot?: Four Ethical and Socio-technical Questions about Wellbeing Robots in the Real World through Community Engagement
- Deep opacity and AI: A threat to XAI and to privacy protection mechanisms
- RUDI: An evidence-based police-centric guide for approaching the development of algorithmic models in policing
- Safe-Control: A Safety Patch for Mitigating Unsafe Content in Text-to-Image Generation Models
- Studying Up Machine Learning Data: Why Talk About Bias When We Mean Power?
- Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting
- We Need to Talk About Data: The Importance of Data Readiness in Natural\n Language Processing
- Impact Assessment Card: Communicating Risks and Benefits of AI Uses
- EmoTale: An Enacted Speech-emotion Dataset in Danish
- Breakable Machine: A K-12 Classroom Game for Transformative AI Literacy Through Spoofing and eXplainable AI (XAI)
- Assessing Trustworthiness of AI Training Dataset using Subjective Logic -- A Use Case on Bias
- An Examination of Fairness of AI Models for Deepfake Detection
- A Word on Machine Ethics: A Response to Jiang et al. (2021)
- Documenting Deployment with Fabric: A Repository of Real-World AI Governance
- Utility is in the Eye of the User: A Critique of NLP Leaderboards
- A Decentralized Approach towards Responsible AI in Social Ecosystems
- KLUE: Korean Language Understanding Evaluation
- OPTIC-ER: A Reinforcement Learning Framework for Real-Time Emergency Response and Equitable Resource Allocation in Underserved African Communities
- Beyond Internal Data: Bounding and Estimating Fairness from Incomplete Data
- Street Review: A Participatory AI-Based Framework for Assessing Streetscape Inclusivity
- TechOps: Technical Documentation Templates for the AI Act
- Component Mismatches Are a Critical Bottleneck to Fielding AI-Enabled Systems in the Public Sector
- Simon Says: Evaluating and Mitigating Bias in Pruned Neural Networks with Knowledge Distillation
- A software engineering perspective on engineering machine learning systems: State of the art and challenges
- Fabricating Holiness: Characterizing Religious Misinformation Circulators on Arabic Social Media
- Towards Experience-Centered AI: A Framework for Integrating Lived Experience in Design and Development
- Dynamic Algorithmic Service Agreements Perspective
- AI and Holistic Review: Informing Human Reading in College Admissions
- If-T: A Benchmark for Type Narrowing
- Dynaword: From One-shot to Continuously Developed Datasets
- Future Illiteracies -- Architectural Epistemology and Artificial Intelligence
- How Growing Toxicity Manifests: A Topic Trajectory Analysis of U.S. Immigration Discourse on Social Media
- CHAIMELEON Project: Creation of a Pan-European Repository of Health Imaging Data for the Development of AI-Powered Cancer Management Tools. [europepmc]
- Reliance on metrics is a fundamental challenge for AI. [europepmc]
- The diagnostic and triage accuracy of digital and online symptom checker tools: a systematic review. [europepmc]
- Possible Bias in Supervised Deep Learning Algorithms for CT Lung Nodule Detection and Classification. [europepmc]
- From compute to care: Lessons learned from deploying an early warning system into clinical practice. [europepmc]
- Piloting a Survey-Based Assessment of Transparency and Trustworthiness with Three Medical AI Tools. [europepmc]
- Clinician's guide to trustworthy and responsible artificial intelligence in cardiovascular imaging. [europepmc]
- Development of an Open-Source Annotated Glaucoma Medication Dataset From Clinical Notes in the Electronic Health Record. [europepmc]
- Survey of Explainable AI Techniques in Healthcare. [europepmc]
- Generalizability challenges of mortality risk prediction models: A retrospective analysis on a multi-center database. [europepmc]
- Open-source tools for behavioral video analysis: Setup, methods, and best practices. [europepmc]
- A Large-scale Synthetic Pathological Dataset for Deep Learning-enabled Segmentation of Breast Cancer. [europepmc]
- Toward fairness in artificial intelligence for medical image analysis: identification and mitigation of potential biases in the roadmap from data collection to model deployment. [europepmc]
- Methodologies for Monitoring Mental Health on Twitter: Systematic Review. [europepmc]
- Large language models encode clinical knowledge. [europepmc]
- Fairness of artificial intelligence in healthcare: review and recommendations. [europepmc]
- Predicting suicide attempts among Norwegian adolescents without using suicide-related items: a machine learning approach. [europepmc]
- A dataset of skin lesion images collected in Argentina for the evaluation of AI tools in this population. [europepmc]
- Use of artificial intelligence in critical care: opportunities and obstacles. [europepmc]
- The METRIC-framework for assessing data quality for trustworthy AI in medicine: a systematic review. [europepmc]
- Applied artificial intelligence for global child health: Addressing biases and barriers. [europepmc]
- Early Prediction of Cardiac Arrest in the Intensive Care Unit Using Explainable Machine Learning: Retrospective Study. [europepmc]
- A toolbox for surfacing health equity harms and biases in large language models. [europepmc]
- FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare. [europepmc]
- AI-driven healthcare: Fairness in AI healthcare: A survey. [europepmc]
Related