2025/07/14 by İsmail Tarım, Tarım, İsmail, Aytuğ Onan +1 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Authorship Attribution and Profiling #Autoregressive model #Computation and Language (cs.CL) #Computational linguistics #FOS: Computer and information sciences #H.3.3 #Hate Speech and Cyberbullying Detection #Hierarchy #I.2.7 #Language model #Metric (unit) #Perplexity #Text Readability and Simplification
paper · pdf · doi:10.48550/arxiv.2507.10475
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/07/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
The rapid advancement of large language models (LLMs) has raised concerns about reliably detecting AI-generated text. Stylometric metrics work well on autoregressive (AR) outputs, but their effectiveness on diffusion-based models is unknown. We present the first systematic comparison of diffusion-generated text (LLaDA) and AR-generated text (LLaMA) using 2 000 samples. Perplexity, burstiness, lexical diversity, readability, and BLEU/ROUGE scores show that LLaDA closely mimics human text in perplexity and burstiness, yielding high false-negative rates for AR-oriented detectors. LLaMA shows much lower perplexity but reduced lexical fidelity. Relying on any single metric fails to separate diffusion outputs from human writing. We highlight the need for diffusion-aware detectors and outline directions such as hybrid models, diffusion-specific stylometric signatures, and robust watermarking.