Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
2024/03/11 by Weixin Liang, Liang, Weixin, Zachary Izzo +22 · 25 voices · 70 citations
Mathematics · Medicine · Psychology · #Artificial Intelligence in Healthcare and Education #Cartography #Computer science #Content (measure theory) #Geography #Law #Mathematics #Peer review #Political science #Psychology #Scale (ratio)
paper · pdf · doi:10.48550/arxiv.2403.07183
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/03/11 · openalex created_date 2024/03/14 · openalex updated_date 2026/07/29
Abstract
We present an approach for estimating the fraction of text in a large corpus which is likely to be substantially modified or produced by a large language model (LLM). Our maximum likelihood model leverages expert-written and AI-generated reference texts to accurately and efficiently examine real-world LLM-use at the corpus level. We apply this approach to a case study of scientific peer review in AI conferences that took place after the release of ChatGPT: ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023. Our results suggest that between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing updates. The circumstances in which generated text occurs offer insight into user behavior: the estimated fraction of LLM-generated text is higher in reviews which report lower confidence, were submitted close to the deadline, and from reviewers who are less likely to respond to author rebuttals. We also observe corpus-level trends in generated text which may be too subtle to detect at the individual level, and discuss the implications of such trends on peer review. We call for future interdisciplinary work to examine how LLM use is changing our information and knowledge practices.
Cited by
- The Aura in the Machine: Genealogy and the Status of the Work of Art in the Generative Era
- Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers?
- AI-Augmented Science and the New Institutional Scarcities
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The Digital Divide in Generative AI: Evidence from Large Language Model Use in College Admissions Essays
- AI use in American newspapers is widespread, uneven, and rarely disclosed
- Can GenAI Improve Academic Performance? Evidence from the Social and Behavioral Sciences
- Prestige over merit: An adapted audit of LLM bias in peer review
- Screening, sorting, and the feedback cycles that imperil peer review
- The Widespread Adoption of Large Language Model-Assisted Writing Across Society
- The Hitchhiker's Guide to Monoculture
- Generative Artificial Intelligence in Scientific Research: Individual Benefits, Collective Risks, and a Framework for Responsible Research with AI
- The Impact Market to Save Conference Peer Review: Decoupling Dissemination and Credentialing
- Evaluating science: A comparison of human and AI reviewers
- CryptoQA: A Large-scale Question-answering Dataset for AI-assisted Cryptography
- Estimating the prevalence of LLM-assisted text in scholarly writing
- Reducing research bureaucracy in UK higher education: Can generative AI assist with the internal evaluation of quality?
- AI-Assisted Writing Is Growing Fastest Among Non-English-Speaking and Less Established Scientists
- Does Scientific Writing Converge to U.S. English? Evidence from Generative AI-Assisted Publications
- Have we reached the beginning of the end for review papers?
- Developing Students’ Statistical Expertise Through Writing in the Age of AI
- The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing
- LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers
- The Rise of Large Language Models and the Direction and Impact of US Federal Research Funding
- ReviewGuard: Enhancing Deficient Peer Review Detection via LLM-Driven Data Augmentation
- Gen-Review: A Large-scale Dataset of AI-Generated (and Human-written) Peer Reviews
- On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
- BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
- No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions
- ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review
- Machines in the Crowd? Measuring the Footprint of Machine-Generated Text on Reddit
- Linguistic Characteristics of AI-Generated Text: A Survey
- Who's Your Judge? On the Detectability of LLM-Generated Judgments
- Shifting norms in scholarly publications: trends in readability, objectivity, authorship, and AI use
- Survivors, Complainers, and Borderliners: Upward Bias in Online Discussions of Academic Conference Reviews
- The Matrix of AI Agency: On the Demarcation Problem in Social Theory
- When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
- The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
- Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
- CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review Detection
- A Multi-Task Evaluation of LLMs' Processing of Academic Text Input
- Echoes of Automation: The Increasing Use of LLMs in Newsmaking
- Lessons from complex systems science for AI governance
- Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and Evaluator
- SIMDAVIS 1.2: phosphonates are outstanding SIM ligands. Crown ethers are not
- Research quality evaluation by AI in the era of Large Language Models: Advantages, disadvantages, and systemic effects
- From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer Review
- The Lock-in Hypothesis: Stagnation by Algorithm
- EMAC+: Embodied Multimodal Agent for Collaborative Planning with VLM+LLM
- GPT Editors, Not Authors: The Stylistic Footprint of LLMs in Academic Preprints
- Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery
- In-Context Watermarks for Large Language Models
- Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
- Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors
- Examining Linguistic Shifts in Academic Writing Before and After the Launch of ChatGPT: A Study on Preprint Papers
- Plagiarism in peer-review reports could be the ‘tip of the iceberg’
- ReVoicer: Conversational Voice Annotation for Human-Centered, LLM-Assisted Peer Review
- “You Cannot Sound Like GPT": Signs of language discrimination and resistance in computer science publishing
- Exploring the change in scientific readability following the release of ChatGPT
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Beyond Via: Analysis and Estimation of the Impact of Large Language Models in Academic Papers
- To Trust or Not to Trust: Authors' Response to AI-based Reviews
- Beyond Detection: Governing GenAI in Academic Peer Review as a Sociotechnical Challenge
- Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift
- Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs
- AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality
- Robust and Fine-Grained Detection of AI Generated Texts
- The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Identifying Aspects in Peer Reviews
Discussions
- A new arXiv preprint from James Zou suggests that quite a few academics are also using ChatGPT to write their peer reviews. I don't have to explain what an utter shitshow this is going to be, do I? [bsky, 371 points, 17 comments]
- Papers are being submitted that are written using AI. Reviews of papers are being written by AI. This is just going to be a circular spiral down the toilet for quality control, isn't it? arxiv.org/abs [bsky, 18 points, 2 comments]
- The optimistic spin on this is that it’s only 10% of the reviews, and that the monitoring techniques developed here are potentially a pretty good way to push that number down. arxiv.org/abs/2403.07183 [bsky, 12 points, 3 comments]
- A lot of our current conversation is situated around the *prevalence* of AI in papers (arxiv.org/abs/2404.01268), reviews (arxiv.org/abs/2403.07183), and so on. But what is, imo, more important is to [bsky, 7 points, 1 comments]
- I've been seeing this egregiously as an AC. I saw a review that had the SAME generic bullet points ("Novelty" etc) listed under pros and cons. It's a serious threat and I don't know how peer review wi [bsky, 4 points, 1 comments]
- AI writing reviews for CS conferences, but not for general science journals (Nature). My hot take: Reviews in CS conferences are so poorly reasoned that using an LLM makes no difference in the quali [bsky, 4 points, 0 comments]
- "Our results suggest that between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs": arxiv.org/pdf/2403.071... [bsky, 4 points, 0 comments]
- This study estimated AI-usage in AI conference reviews not by trying to "detect" AI-written documents but by estimating the usage at the level of the whole corpus, with a convincing validation. Not [bsky, 3 points, 0 comments]
- A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews [hn, 2 points, 0 comments]
- Interesting preprint examining things like adjective frequency in peer review text, noting dramatic shifts in top adjectives produced disproportionately by AI. I.e., peer reviewers are using ChatGPT a [bsky, 1 points, 0 comments]
- A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews [hn, 1 points, 0 comments]
- via Carl T. Bergstrom: A new arXiv preprint from James Zou suggests that quite a few academics are also using ChatGPT to write their peer reviews. I don't have to explain what an utter shitshow this [bsky, 1 points, 1 comments]
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews https://arxiv.org/abs/2403.07183 [bsky, 1 points, 0 comments]
- That's the whole problem. People are busy and/or lazy, so they glance at ChatGPT output and go "yeah, good enough" even when it isn't. There's a disturbing trend for scientific peer review to be writ [bsky, 1 points, 1 comments]
- What about the presence of ‘commendable’ ‘intricate’ ‘notable’ ‘innovative’ etc ?… because: arxiv.org/abs/2403.07183. Indeed, maybe the LLM is suggesting ‘nuance’ as a fit for sociology because of al [bsky, 1 points, 1 comments]
- #icml2024 paper: how are LLMs used in reviews? 10% of ICLR sentences are auto-generated. More LLM usage when submitting later Less when referring to at least one other paper arxiv.org/abs/2403.07183 � [bsky, 1 points, 0 comments]
- I find the robustness check that the authors do on this in sections 4.5 and 4.6 not entirely convincing. arxiv.org/abs/2403.07183 [bsky, 1 points, 1 comments]
- arxiv.org/abs/2403.07183 [bsky, 1 points, 0 comments]
- Thinking again about how OpenAI said in federal court that no user would believe its false outputs are real arxiv.org/abs/2403.07183 [bsky, 1 points, 0 comments]
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews arxiv.org/pdf/2403.071... [bsky, 1 points, 0 comments]
- Monitoring AI-Modified Content: Impact of ChatGPT on AI Conference Peer Reviews [hn, 1 points, 0 comments]
- 🧪 🤖 #AcademicSky Direct link to the preprint: arxiv.org/abs/2403.07183 [bsky, 0 points, 0 comments]
- or they're cringe from being LinkedIn-coded, though they seem pretty on-brand for academic papers [bsky, 0 points, 1 comments]
- I've been torturing my dad by sending him articles like these. (Clickable link for those who are interested: arxiv.org/abs/2403.07183 ) [bsky, 0 points, 0 comments]
- "Our results suggest that between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing u [bsky, 0 points, 0 comments]
Related