From Reflection to Repair: A Scoping Review of Dataset Documentation Tools
2026/02/17 by Pedro Reynolds-Cuéllar, Marisol Wong-Villacres, Adriana Alvarado Garcia +1 · 1 voice
Computer Science · #cs.SE #cs.AI #cs.CY #cs.HC
paper · pdf · doi:10.1145/3772318.3791344
Abstract
Dataset documentation is widely recognized as essential for the responsible development of automated systems. Despite growing efforts to support documentation through different kinds of artifacts, little is known about the motivations shaping documentation tool design or the factors hindering their adoption. We present a systematic review supported by mixed-methods analysis of 59 dataset documentation publications to examine the motivations behind building documentation tools, how authors conceptualize documentation practices, and how these tools connect to existing systems, regulations, and cultural norms. Our analysis shows four persistent patterns in dataset documentation conceptualization that potentially impede adoption and standardization: unclear operationalizations of documentation's value, decontextualized designs, unaddressed labor demands, and a tendency to treat integration as future work. Building on these findings, we propose a shift in Responsible AI tool design toward institutional rather than individual solutions, and outline actions the HCI community can take to enable sustainable documentation practices.
Citations
- A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset
- Position: The Most Expensive Part of an LLM should be its Training Data
- Datasheets for AI and medical datasets (DAIMS): a data validation and documentation framework before machine learning analysis in medical research
- Datasheets for Healthcare AI: A Framework for Transparency and Bias Mitigation
- The State of Data Curation at NeurIPS: An Assessment of Dataset Development Practices in the Datasets and Benchmarks Track
- Improving governance outcomes through AI documentation: Bridging theory and practice
- A Standardized Machine-readable Dataset Documentation Format for Responsible AI
- U Can't Gen This? A Survey of Intellectual Property Protection Methods for Data in Generative AI
- Using Large Language Models to Enrich the Documentation of Datasets for Machine Learning
- Croissant: A Metadata Format for ML-Ready Datasets
- Guidelines for Integrating Value Sensitive Design in Responsible AI Toolkits
- On the Challenges and Opportunities in Generative AI
- Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on Hugging Face
- Open Datasheets: Machine-readable Documentation for Open Datasets and Responsible AI Assessments
- ‘What they’re not telling you about ChatGPT’: exploring the discourse of AI in UK news media headlines
- On Responsible Machine Learning Datasets with Fairness, Privacy, and Regulatory Norms
- The Foundation Model Transparency Index
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Investigating Practices and Opportunities for Cross-functional Collaboration around AI Fairness in Industry Practice
- Walking the Walk of AI Ethics: Organizational Challenges and the Individualization of Risk among Ethics Entrepreneurs
- Datasheet for Subjective and Objective Quality Assessment Datasets
- Right the docs: Characterising voice dataset documentation practices used in machine learning
- GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
- Network Report: A Structured Description for Network Datasets
- Understanding Machine Learning Practitioners' Data Documentation Perceptions, Needs, Challenges, and Desiderata
- On the Origin of Hallucinations in Conversational Models: Is it the Datasets or the Models?
- Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI
- Healthsheet: Development of a Transparency Artifact for Health Datasets
- Can Machines Help Us Answering Question 16 in Datasheets, and In Turn Reflecting on Inappropriate Content?
- The Dataset Nutrition Label (2nd Gen): Leveraging Context to Mitigate Harms in Artificial Intelligence
- Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for Support
- Mitigating Dataset Harms Requires Stewardship: Lessons from 1000 Papers
- Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus
- What's in the Box? A Preliminary Analysis of Undesirable Content in the Common Crawl Corpus
- On the Dangers of Stochastic Parrots
- Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure
- Data Readiness Report
- Continued post-retraction citation of a fraudulent clinical trial report, 11 years after it was retracted for falsifying data
- Ensuring Dataset Quality for Machine Learning Certification
- One size fits all? What counts as quality practice in (reflexive) thematic analysis?
- A Methodology for Creating AI FactSheets
- Where Responsible AI meets Reality: Practitioner Perspectives on Enablers for shifting Organizational Practices
- FactSheets: Increasing trust in AI services through supplier's declarations of conformity
- Improving fairness in machine learning systems: What do industry practitioners need?
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation
- The Dataset Nutrition Label: A Framework To Drive Higher Data Quality Standards
- No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World
- Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings
Discussions
Related