A Comprehensive Dataset for Human vs. AI Generated Text Detection
2025/10/26 by Roy, Rajarshi, Imanpour, Nasrin, Aziz, Ashhar +17
#Computation and Language (cs.CL) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2510.22874
Abstract
The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authenticity, misinformation, and trustworthiness. Addressing the challenge of reliably detecting AI-generated text and attributing it to specific models requires large-scale, diverse, and well-annotated datasets. In this work, we present a comprehensive dataset comprising over 58,000 text samples that combine authentic New York Times articles with synthetic versions generated by multiple state-of-the-art LLMs including Gemma-2-9b, Mistral-7B, Qwen-2-72B, LLaMA-8B, Yi-Large, and GPT-4-o. The dataset provides original article abstracts as prompts, full human-authored narratives. We establish baseline results for two key tasks: distinguishing human-written from AI-generated text, achieving an accuracy of 58.35%, and attributing AI texts to their generating models with an accuracy of 8.92%. By bridging real-world journalistic content with modern generative models, the dataset aims to catalyze the development of robust detection and attribution methods, fostering trust and transparency in the era of generative AI. Our dataset is available at: https://huggingface.co/datasets/gsingh1-py/train.
Citations
- Learning from Generalization Patterns: An Evaluation-Driven Approach to Enhanced Data Augmentation for Fine-Tuning Small Language Models
- Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID
- FAID: Fine-Grained AI-Generated Text Detection Using Multi-Task Auxiliary and Multi-Level Contrastive Learning
- DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
- Watermarking Large Language Models and the Generated Content: Opportunities and Challenges
- MMCFND: Multimodal Multilingual Caption-aware Fake News Detection for Low-resource Indic Languages
- Overview of Factify5WQA: Fact Verification through 5W Question-Answering
- Scalable watermarking for identifying large language model outputs
- Evidence-backed Fact Checking using RAG and Few-Shot In-Context Learning with LLMs
- LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection
- Gemma 2: Improving Open Language Models at a Practical Size
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
- Enhancing Low-Resource LLMs Classification with PEFT and Synthetic Data
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
- Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature
- Synthetic Data Generation in Low-Resource Settings via Fine-Tuning of Large Language Models
- Findings of Factify 2: Multimodal Fake News Detection
- RADAR: Robust AI-Text Detection via Adversarial Learning
- M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection
- Factify 2: A Multimodal Fake News and Satire News Dataset
- Cognitive Constraint Simulation and the Geometry of Human Authorship: A First-Principles Theory of AI Text Detection
- How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection
- TURINGBENCH: A Benchmark Environment for Turing Test in the Age of\n Neural Text Generation
- COVID-19 Tests Gone Rogue: Privacy, Efficacy, Mismanagement and Misunderstandings
- Graph Neural Networks with Continual Learning for Fake News Detection from Social Media
- Language Models are Few-Shot Learners
- Rumor Detection on Social Media with Bi-Directional Graph Convolutional Networks
- Textual misinformation on Reddit
- CTRL: A Conditional Transformer Language Model for Controllable Generation
- MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact\n Checking of Claims
- GEAR: Graph-based Evidence Aggregating and Reasoning for Fact Verification
- XLNet: Generalized Autoregressive Pretraining for Language Understanding
- Defending Against Neural Fake News
- Cross-lingual Language Model Pretraining
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- FakeNewsNet: A Data Repository with News Content, Social Context and Spatialtemporal Information for Studying Fake News on Social Media
- FakeNewsNet: A Data Repository with News Content, Social Context, and Spatiotemporal Information for Studying Fake News on Social Media
- FEVER: a large-scale dataset for Fact Extraction and VERification
- "Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News\n Detection
Related