vix.ing · top · new · best · stats

WaterDrum: Watermarking for Data-centric Unlearning Metric

2025/05/08 by Xinyang Lu, Lu, Xinyang, Xinyuan Niu +16 · 4 citations
Computer Science · #Advanced Graph Neural Networks #Adversarial Robustness in Machine Learning #Benchmark (surveying) #Code (set theory) #Digital watermarking #Exploit #Generative Adversarial Networks and Image Synthesis #Metric (unit) #Retraining #Set (abstract data type)

paper · pdf · doi:10.48550/arxiv.2505.05064

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2025/05/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Large language model (LLM) unlearning is critical in real-world applications where it is necessary to efficiently remove the influence of private, copyrighted, or harmful data from some users. Existing utility-centric unlearning metrics (based on model utility) may fail to accurately evaluate the extent of unlearning in realistic settings such as when the forget and retain sets have semantically similar content and/or retraining the model from scratch on the retain set is impractical. This paper presents the first data-centric unlearning metric for LLMs called WaterDrum that exploits robust text watermarking to overcome these limitations. We introduce new benchmark datasets (with different levels of data similarity) for LLM unlearning that can be used to rigorously evaluate unlearning algorithms via WaterDrum. Our code is available at https://github.com/lululu008/WaterDrum and our new benchmark datasets are released at https://huggingface.co/datasets/Glow-AI/WaterDrum-Ax.

Citations

Cited by

Related