vix.ing · top · new · best · stats · spec

DIS-CO: Discovering Copyrighted Content in VLMs Training Data

2025/02/24 by André V. Duarte, Xuandong Zhao, Duarte, André V. +5 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · #Artificial Intelligence (cs.AI) #Biomedical Text Mining and Ontologies #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #I.2 #Imbalanced Data Classification Techniques #Machine Learning (cs.LG) #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2502.17358

openalex publication_date 2025/02/24 · openalex created_date 2025/10/15 · openalex updated_date 2026/07/28

Abstract

How can we verify whether copyrighted content was used to train a large vision-language model (VLM) without direct access to its training data? Motivated by the hypothesis that a VLM is able to recognize images from its training corpus, we propose DIS-CO, a novel approach to infer the inclusion of copyrighted content during the model's development. By repeatedly querying a VLM with specific frames from targeted copyrighted material, DIS-CO extracts the content's identity through free-form text completions. To assess its effectiveness, we introduce MovieTection, a benchmark comprising 14,000 frames paired with detailed captions, drawn from films released both before and after a model's training cutoff. Our results show that DIS-CO significantly improves detection performance, nearly doubling the average AUC of the best prior method on models with logits available. Our findings also highlight a broader concern: all tested models appear to have been exposed to some extent to copyrighted content. Our code and data are available at https://github.com/avduarte333/DIS-CO

Cited by

Related