vix.ing · top · new · best · stats · spec

CATMark: A Context-Aware Thresholding Framework for Robust Cross-Task Watermarking in Large Language Models

2025/09/27 by Zhang, Yu, Liu, Shuliang, Yang, Xu +1
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences

paper · doi:10.48550/arxiv.2510.02342

Abstract

Watermarking algorithms for Large Language Models (LLMs) effectively identify machine-generated content by embedding and detecting hidden statistical features in text. However, such embedding leads to a decline in text quality, especially in low-entropy scenarios where performance needs improvement. Existing methods that rely on entropy thresholds often require significant computational resources for tuning and demonstrate poor adaptability to unknown or cross-task generation scenarios. We propose Context-Aware Threshold watermarking (\myalgo), a novel framework that dynamically adjusts watermarking intensity based on real-time semantic context. \myalgo partitions text generation into semantic states using logits clustering, establishing context-aware entropy thresholds that preserve fidelity in structured content while embedding robust watermarks. Crucially, it requires no pre-defined thresholds or task-specific tuning. Experiments show \myalgo improves text quality in cross-tasks without sacrificing detection accuracy.

Citations

Related