AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
2025/09/29 by Zhou, Hongyi, Zhu, Jin, Su, Pingfan +4 · 1 citation
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
paper · doi:10.48550/arxiv.2510.01268
Abstract
We study the problem of determining whether a piece of text has been authored by a human or by a large language model (LLM). Existing state of the art logits-based detectors make use of statistics derived from the log-probability of the observed text evaluated using the distribution function of a given source LLM. However, relying solely on log probabilities can be sub-optimal. In response, we introduce AdaDetectGPT -- a novel classifier that adaptively learns a witness function from training data to enhance the performance of logits-based detectors. We provide statistical guarantees on its true positive rate, false positive rate, true negative rate and false negative rate. Extensive numerical studies show AdaDetectGPT nearly uniformly improves the state-of-the-art method in various combination of datasets and LLMs, and the improvement can reach up to 37%. A python implementation of our method is available at https://github.com/Mamba413/AdaDetectGPT.
Citations
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- An Overview of Large Language Models for Statisticians
- Qwen2.5-VL Technical Report
- A Statistical Hypothesis Testing Framework for Data Misappropriation Detection in Large Language Models
- Qwen2.5 Technical Report
- Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection
- Robust Detection of Watermarks for Large Language Models Under Human Edits
- DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive Learning
- GPT-4o System Card
- A Watermark for Order-Agnostic Language Models
- Scalable watermarking for identifying large language model outputs
- The Llama 3 Herd of Models
- Gemma 2: Improving Open Language Models at a Practical Size
- DALD: Improving Logits-based Detector without Logits from Black-box LLMs
- Edit Distance Robust Watermarks via Indexing Pseudorandom Codes
- ReMoDetect: Reward Models Recognize Aligned LLM's Generations
- A Statistical Framework of Watermarks for Large Language Models: Pivot, Detection Efficiency and Optimal Rules
- WaterMax: breaking the LLM watermark detectability-robustness-quality trade-off
- Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language Models
- Detecting Machine-Generated Texts by Multi-Population Aware Optimization for Maximum Mean Discrepancy
- Adaptive Text Watermark for Large Language Models
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Optimizing watermarks for large language models
- A Simple yet Efficient Ensemble Approach for AI-generated Text Detection
- A Survey on Detection of LLMs-Generated Content
- A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
- A Resilient and Accessible Distribution-Preserving Watermark for Large Language Models
- Mistral 7B
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
- Unbiased Watermark for Large Language Models
- Robust Distortion-free Watermarks for Language Models
- OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
- RADAR: Robust AI-Text Detection via Adversarial Learning
- Provable Robust Watermarking for AI-Generated Text
- Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts
- Multiscale Positive-Unlabeled Detection of AI-Generated Texts
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text
- Undetectable Watermarks for Language Models
- Ghostbuster: Detecting Text Ghostwritten by Large Language Models
- DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text
- DPIC: Decoupling Prompt and Intrinsic Characteristics for LLM Generated Text Detection
- GPT detectors are biased against non-native English writers
- Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense
- Stylometric Detection of AI-Generated Text in Twitter Timelines
- ChatGPT or Human? Detect and Explain. Explaining Decisions of Machine Learning Model for Detecting Short ChatGPT-generated Text
- Cognitive Constraint Simulation and the Geometry of Human Authorship: A First-Principles Theory of AI Text Detection
- A Watermark for Large Language Models
- How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection
- PaLM: Scaling Language Modeling with Pathways
- Detecting Fake News Using Machine Learning : A Systematic Literature Review
- Statistical Inference of the Value Function for Reinforcement Learning in Infinite Horizon Settings
- Automatic Detection of Generated Text is Easiest when Humans are Fooled
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Release Strategies and the Social Impacts of Language Models
- GLTR: Statistical Detection and Visualization of Generated Text
- Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional\n Neural Networks for Extreme Summarization
- Hierarchical Neural Story Generation
- SQuAD: 100,000+ Questions for Machine Comprehension of Text
- Character-level Convolutional Networks for Text Classification
- Fast learning rates for plug-in classifiers under the margin condition
Cited by
Related