Multimodal datasets: misogyny, pornography, and malignant stereotypes
2021/10/05 by Abeba Birhane, Vinay Uday Prabhu, Birhane, Abeba +3 · 12 voices · 49 citations
Social Sciences · Psychology · Computer Science · #Gender, Feminism, and Media #Sexuality, Behavior, and Technology #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2110.01963
Abstract
We have now entered the era of trillion parameter machine learning models trained on billion-sized datasets scraped from the internet. The rise of these gargantuan datasets has given rise to formidable bodies of critical work that has called for caution while generating these large datasets. These address concerns surrounding the dubious curation practices used to generate these datasets, the sordid quality of alt-text data available on the world wide web, the problematic content of the CommonCrawl dataset often used as a source for training large language models, and the entrenched biases in large-scale visio-linguistic models (such as OpenAI's CLIP model) trained on opaque datasets (WebImageText). In the backdrop of these specific calls of caution, we examine the recently released LAION-400M dataset, which is a CLIP-filtered dataset of Image-Alt-text pairs parsed from the Common-Crawl dataset. We found that the dataset contains, troublesome and explicit images and text pairs of rape, pornography, malign stereotypes, racist and ethnic slurs, and other extremely problematic content. We outline numerous implications, concerns and downstream harms regarding the current state of large scale datasets while raising open questions for various stakeholders including the AI community, regulators, policy makers and data subjects.
Citations
Cited by
Discussions
- once you learn how much training data images of women come from porn, it illuminates so much. this is more than bad to look at, it's the rot of misogyny in full display. arxiv.org/pdf/2110.01963 [bsky, 2118 points, 15 comments]
- Pour souvenir, ou pour info : Concernant les corps féminins, énormément de modèle sont entrainées sur de la donnée venant du porno. Pensez-y le visage d'une femme est "amélioré", et bien maquillé. Enc [bsky, 56 points, 3 comments]
- no. feel welcome to read my work on dataset audit 1. arxiv.org/pdf/2110.01963 2. ieeexplore.ieee.org/abstract/doc... 3. proceedings.neurips.cc/paper_files/... 4. dl.acm.org/doi/pdf/10.1... [bsky, 45 points, 2 comments]
- It's both though. AI image generators have clearly been trained on social media content, but it also has been found that some data sets used to train generative AI contain pornography and CSAM in thei [bsky, 10 points, 1 comments]
- Correct. arxiv.org/pdf/2110.01963 [bsky, 6 points, 0 comments]
- This is a pattern of behaviour for LAION Contains “rape, pornography, malign stereotypes, racist and ethnic slurs, and other extremely problematic content” arxiv.org/abs/2110.01963 “LAION scraped m [bsky, 6 points, 0 comments]
- people blew off the research showing that these algorithms were trained on images of sexual violence...! arxiv.org/abs/2110.01963 [bsky, 4 points, 0 comments]
- Misogyny, pornography, and malignant stereotypes in LAION-400M image dataset [hn, 2 points, 0 comments]
- I'll never understand how this paper exists, yet somehow 4 full years later we still ended up here. At no point did any leadership in the industry read this and ask: maybe unscalable and unspeakable h [bsky, 1 points, 0 comments]
- On the same topic, there is a paper from 2021 titled “Multimodal datasets: misogyny, pornography, and malignant stereotypes”. arxiv.org/pdf/2110.019... [bsky, 0 points, 0 comments]
- Yes. I read things like this. If only someone had warned.... arxiv.org/abs/2110.01963 [bsky, 0 points, 0 comments]
- Een belangrijk probleem bij dit soort onderzoek is historisch, culturele en sociale bias. De AI is getraind op miljoenen beeld-tekst combinaties van het internet. Hedendaagse stereotype beïnvloeden hi [bsky, 0 points, 1 comments]
Related