People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text
2025/01/26 by Jenna Russell, Marzena Karpinska, Russell, Jenna +3 · 30 voices · 17 citations
#cs.CL #cs.AI
paper · pdf · doi:10.48550/arxiv.2501.15654
Abstract
In this paper, we study how well humans can detect text generated by commercial LLMs (GPT-4o, Claude, o1). We hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions. Our experiments show that annotators who frequently use LLMs for writing tasks excel at detecting AI-generated text, even without any specialized training or feedback. In fact, the majority vote among five such "expert" annotators misclassifies only 1 of 300 articles, significantly outperforming most commercial and open-source detectors we evaluated even in the presence of evasion tactics like paraphrasing and humanization. Qualitative analysis of the experts' free-form explanations shows that while they rely heavily on specific lexical clues ('AI vocabulary'), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity) that are challenging to assess for automatic detectors. We release our annotated dataset and code to spur future research into both human and automated detection of AI-generated text.
Citations
Cited by
Discussions
- One of the great ironies of AI writing is that the only people who can detect it with accuracy are people who use AI for writing a lot (at least if you take a majority vote among five such people) Non [bsky, 97 points, 3 comments]
- People who frequently use ChatGPT for writing tasks can detect AI-generated text [hn, 12 points, 7 comments]
- Frequent ChatGPT users are accurate detectors of AI-generated text (2025) [hn, 11 points, 2 comments]
- If this finding holds, it is (perhaps unfortunately) reason for teachers to get familiar with LLMs for text generation arxiv.org/abs/2501.15654 [bsky, 11 points, 0 comments]
- Last year, only one AI program was as good as trained humans. arxiv.org/abs/2501.15654 They're both really good though, well over 99% [bsky, 7 points, 1 comments]
- I don't normally repost images without alt text, but all the images in the thread are taken from the research paper linked further down, and it is a very interesting paper PDF: arxiv.org/pdf/2501.1565 [bsky, 6 points, 0 comments]
- The honest tension: Not quantifiable, not statistically provable, but a real gut feeling—more than vibes—that anyone who's used these tools quietly recognizes. arxiv.org/abs/2501.15654 [bsky, 6 points, 0 comments]
- Research Article (Preprint): People Who Frequently Use #ChatGPT For #Writing Tasks are Accurate and Robust Detectors of AI-Generated Text arxiv.org/abs/2501.15654 #AI [bsky, 5 points, 0 comments]
- Les gens qui ont l'habitude d'utiliser ChatGPT pour rédiger détectent assez bien quand un texte en est issu, même s'il a été retravaillé. arxiv.org/abs/2501.15654 Est-ce que ce serait une bonne idée [bsky, 4 points, 0 comments]
- one illustrative example: “annotators who frequently use LLMs for writing tasks excel at detecting AI-generated text, even without any specialized training or feedback” arxiv.org/abs/2501.15654 [bsky, 4 points, 1 comments]
- Can academics please stop regurgitating the myth that we can't detect genAI. '[T]the majority vote among five such “expert” annotators misclassifies only 1 of 300 [genAI written] articles [...] even i [bsky, 3 points, 1 comments]
- arxiv.org/pdf/2501.15654 is the data. [bsky, 3 points, 1 comments]
- A new study suggests that people who use AI for writing are more able to detect AI writing than automated scanner tools. My current LLM pet peeve is how they use language like load-bearing, structural [bsky, 3 points, 3 comments]
- People who use ChatGPT for writing are accurate detectors of AI text (2025) [hn, 3 points, 0 comments]
- Изглежда че единствения начин да се засичат писания творени от ИИ за момента, е хора да ги ползват и да трупат опит. Human mind is yet to me matched by machines 👌🏻 [bsky, 2 points, 1 comments]
- at least one study suggests that it actually can detect several kinds of adversarial model quite accurately, as can people who are familiar with ai writing arxiv.org/abs/2501.156... [bsky, 2 points, 1 comments]
- Link found in last post of thread 😀 (but putting it here again) arxiv.org/abs/2501.15654 [bsky, 2 points, 1 comments]
- pretty much yeah arxiv.org/pdf/2501.15654 [bsky, 2 points, 0 comments]
- People who use ChatGPT for writing are robust detectors of AI-generated text [hn, 2 points, 0 comments]
- Oh dear the converse will be true wont it? I’ve already been called an LLM from my natural speech. The same people didn’t recognize Claude or Gemini posted verbatim. Oh dear arxiv.org/abs/2501.156... [bsky, 1 points, 0 comments]
- It’s not “just vibes.” A person can reliably detect ai with enough exposure. arxiv.org/abs/2501.15654 [bsky, 1 points, 0 comments]
- Likewise, "People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text" arxiv.org/abs/2501.156... Are these people, too, having a hard time paying attent [bsky, 1 points, 1 comments]
- People who use ChatGPT for writing are accurate detectors of AI-generated text [hn, 1 points, 0 comments]
- Frequent users of ChatGPT are robust detectors of AI text [hn, 1 points, 2 comments]
- I haven't reviewed all the studies on this, but I recall this one arxiv.org/abs/2501.15654 for such a list, it would be less about the average person's ability to recognize, and more about someone who [bsky, 0 points, 0 comments]
- Quins recursos tenim? Doncs el que heu comentat: Els docents que utilitzen la IAG són els detectors més fiables. En aquest article, revisors que utilitzen freqüentment la IAG van demostrar una habilit [bsky, 0 points, 0 comments]
- 人間ってすげーなー 正直、人間にはAIのエッセイの検出できっこないと思ってた 検出率100%(300中299正解)か、人間なめてたわ arxiv.org/abs/2501.156... まぁ、自分には無理だな [bsky, 0 points, 0 comments]
- First responders build up cognitive responses while wading through the graphite of #SyntheticChernobyl. https://arxiv.org/abs/2501.15654 [bsky, 0 points, 1 comments]
- I think we need to be really cautious about using AI detectors in this fashion. First, independent evaluations have suggested the false positive rate for these tools (including Pangram) is higher than [bsky, 0 points, 1 comments]
- Why those that use AI to write are far better at spotting AI writing than those who don't use AI to write. arxiv.org/pdf/2501.15654 [bsky, 0 points, 0 comments]
Related