Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
2024/03/11 by Weixin Liang, Liang, Weixin, Zachary Izzo +22 · 25 voices · 43 citations
Medicine · #Artificial Intelligence in Healthcare and Education
paper · pdf · doi:10.48550/arxiv.2403.07183
Abstract
We present an approach for estimating the fraction of text in a large corpus which is likely to be substantially modified or produced by a large language model (LLM). Our maximum likelihood model leverages expert-written and AI-generated reference texts to accurately and efficiently examine real-world LLM-use at the corpus level. We apply this approach to a case study of scientific peer review in AI conferences that took place after the release of ChatGPT: ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023. Our results suggest that between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing updates. The circumstances in which generated text occurs offer insight into user behavior: the estimated fraction of LLM-generated text is higher in reviews which report lower confidence, were submitted close to the deadline, and from reviewers who are less likely to respond to author rebuttals. We also observe corpus-level trends in generated text which may be too subtle to detect at the individual level, and discuss the implications of such trends on peer review. We call for future interdisciplinary work to examine how LLM use is changing our information and knowledge practices.
Cited by
- The Aura in the Machine: Genealogy and the Status of the Work of Art in the Generative Era
- Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers?
- AI-Augmented Science and the New Institutional Scarcities
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The Digital Divide in Generative AI: Evidence from Large Language Model Use in College Admissions Essays
- AI use in American newspapers is widespread, uneven, and rarely disclosed
- Can GenAI Improve Academic Performance? Evidence from the Social and Behavioral Sciences
- Prestige over merit: An adapted audit of LLM bias in peer review
- Screening, sorting, and the feedback cycles that imperil peer review
- The Widespread Adoption of Large Language Model-Assisted Writing Across Society
- The Hitchhiker's Guide to Monoculture
- Generative Artificial Intelligence in Scientific Research: Individual Benefits, Collective Risks, and a Framework for Responsible Research with AI
- The Impact Market to Save Conference Peer Review: Decoupling Dissemination and Credentialing
- Evaluating science: A comparison of human and AI reviewers
- CryptoQA: A Large-scale Question-answering Dataset for AI-assisted Cryptography
- Estimating the prevalence of LLM-assisted text in scholarly writing
- Reducing research bureaucracy in UK higher education: Can generative AI assist with the internal evaluation of quality?
- AI-Assisted Writing Is Growing Fastest Among Non-English-Speaking and Less Established Scientists
- Generative AI as a Linguistic Equalizer in Global Science
- Have we reached the beginning of the end for review papers?
- Developing Students’ Statistical Expertise Through Writing in the Age of AI
- The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing
- LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers
- The Rise of Large Language Models and the Direction and Impact of US Federal Research Funding
- ReviewGuard: Enhancing Deficient Peer Review Detection via LLM-Driven Data Augmentation
- Gen-Review: A Large-scale Dataset of AI-Generated (and Human-written) Peer Reviews
- On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
- BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
- No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions
- ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review
- Machines in the Crowd? Measuring the Footprint of Machine-Generated Text on Reddit
- Linguistic Characteristics of AI-Generated Text: A Survey
- Who's Your Judge? On the Detectability of LLM-Generated Judgments
- Shifting norms in scholarly publications: trends in readability, objectivity, authorship, and AI use
- Survivors, Complainers, and Borderliners: Upward Bias in Online Discussions of Academic Conference Reviews
- The Matrix of AI Agency: On the Demarcation Problem in Social Theory
- When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
- The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
- Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
- CoCoNUTS: Concentrating on Content while Neglecting Uninformative Textual Styles for AI-Generated Peer Review Detection
- A Multi-Task Evaluation of LLMs' Processing of Academic Text Input
- Echoes of Automation: The Increasing Use of LLMs in Newsmaking
- Lessons from complex systems science for AI governance
Discussions
- A new arXiv preprint from James Zou suggests that quite a few academics are also using ChatGPT to write their peer reviews. I don't have to explain what an utter shitshow this is going to be, do I? [bsky, 371 points, 17 comments]
- Papers are being submitted that are written using AI. Reviews of papers are being written by AI. This is just going to be a circular spiral down the toilet for quality control, isn't it? arxiv.org/abs [bsky, 18 points, 2 comments]
- The optimistic spin on this is that it’s only 10% of the reviews, and that the monitoring techniques developed here are potentially a pretty good way to push that number down. arxiv.org/abs/2403.07183 [bsky, 12 points, 3 comments]
- A lot of our current conversation is situated around the *prevalence* of AI in papers (arxiv.org/abs/2404.01268), reviews (arxiv.org/abs/2403.07183), and so on. But what is, imo, more important is to [bsky, 7 points, 1 comments]
- I've been seeing this egregiously as an AC. I saw a review that had the SAME generic bullet points ("Novelty" etc) listed under pros and cons. It's a serious threat and I don't know how peer review wi [bsky, 4 points, 1 comments]
- AI writing reviews for CS conferences, but not for general science journals (Nature). My hot take: Reviews in CS conferences are so poorly reasoned that using an LLM makes no difference in the quali [bsky, 4 points, 0 comments]
- "Our results suggest that between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs": arxiv.org/pdf/2403.071... [bsky, 4 points, 0 comments]
- This study estimated AI-usage in AI conference reviews not by trying to "detect" AI-written documents but by estimating the usage at the level of the whole corpus, with a convincing validation. Not [bsky, 3 points, 0 comments]
- A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews [hn, 2 points, 0 comments]
- Interesting preprint examining things like adjective frequency in peer review text, noting dramatic shifts in top adjectives produced disproportionately by AI. I.e., peer reviewers are using ChatGPT a [bsky, 1 points, 0 comments]
- A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews [hn, 1 points, 0 comments]
- via Carl T. Bergstrom: A new arXiv preprint from James Zou suggests that quite a few academics are also using ChatGPT to write their peer reviews. I don't have to explain what an utter shitshow this [bsky, 1 points, 1 comments]
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews https://arxiv.org/abs/2403.07183 [bsky, 1 points, 0 comments]
- That's the whole problem. People are busy and/or lazy, so they glance at ChatGPT output and go "yeah, good enough" even when it isn't. There's a disturbing trend for scientific peer review to be writ [bsky, 1 points, 1 comments]
- What about the presence of ‘commendable’ ‘intricate’ ‘notable’ ‘innovative’ etc ?… because: arxiv.org/abs/2403.07183. Indeed, maybe the LLM is suggesting ‘nuance’ as a fit for sociology because of al [bsky, 1 points, 1 comments]
- #icml2024 paper: how are LLMs used in reviews? 10% of ICLR sentences are auto-generated. More LLM usage when submitting later Less when referring to at least one other paper arxiv.org/abs/2403.07183 � [bsky, 1 points, 0 comments]
- I find the robustness check that the authors do on this in sections 4.5 and 4.6 not entirely convincing. arxiv.org/abs/2403.07183 [bsky, 1 points, 1 comments]
- arxiv.org/abs/2403.07183 [bsky, 1 points, 0 comments]
- Thinking again about how OpenAI said in federal court that no user would believe its false outputs are real arxiv.org/abs/2403.07183 [bsky, 1 points, 0 comments]
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews arxiv.org/pdf/2403.071... [bsky, 1 points, 0 comments]
- Monitoring AI-Modified Content: Impact of ChatGPT on AI Conference Peer Reviews [hn, 1 points, 0 comments]
- 🧪 🤖 #AcademicSky Direct link to the preprint: arxiv.org/abs/2403.07183 [bsky, 0 points, 0 comments]
- or they're cringe from being LinkedIn-coded, though they seem pretty on-brand for academic papers [bsky, 0 points, 1 comments]
- I've been torturing my dad by sending him articles like these. (Clickable link for those who are interested: arxiv.org/abs/2403.07183 ) [bsky, 0 points, 0 comments]
- "Our results suggest that between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing u [bsky, 0 points, 0 comments]
Related