Delving into LLM-assisted writing in biomedical publications through excess vocabulary
2024/06/11 by Dmitry Kobak, Rita González-Márquez, Kobak, Dmitry +5 · 24 voices · 32 citations
Medicine · Computer Science · #Artificial Intelligence in Healthcare and Education #Topic Modeling #Machine Learning in Healthcare
paper · pdf · doi:10.1126/sciadv.adt3813
Abstract
Large language models (LLMs) like ChatGPT can generate and revise text with human-level performance. These models come with clear limitations, can produce inaccurate information, and reinforce existing biases. Yet, many scientists use them for their scholarly writing. But how widespread is such LLM usage in the academic literature? To answer this question for the field of biomedical research, we present an unbiased, large-scale approach: We study vocabulary changes in more than 15 million biomedical abstracts from 2010 to 2024 indexed by PubMed and show how the appearance of LLMs led to an abrupt increase in the frequency of certain style words. This excess word analysis suggests that at least 13.5% of 2024 abstracts were processed with LLMs. This lower bound differed across disciplines, countries, and journals, reaching 40% for some subcorpora. We show that LLMs have had an unprecedented impact on scientific writing in biomedical research, surpassing the effect of major world events such as the COVID pandemic.
Citations
Cited by
- Will the widespread use of large language models in scientific writing undermine scientists’ critical thinking?
- StoryScope: Investigating idiosyncrasies in AI fiction
- How AI use in scholarly publishing threatens research integrity, lessens trust, and invites misinformation
- Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High-Quality Books
- Scientific production in the era of large language models
- Surface Reading LLMs: Synthetic Text and its Styles
- AI use in American newspapers is widespread, uneven, and rarely disclosed
- Can GenAI Improve Academic Performance? Evidence from the Social and Behavioral Sciences
- Towards Automating Scientific Review with Google's Paper Assistant Tool
- Language machines: Toward a linguistic anthropology of large language models
- Generative AI and the future of scientometrics: current topics and future questions
- Generative Artificial Intelligence in Scientific Research: Individual Benefits, Collective Risks, and a Framework for Responsible Research with AI
- General-purpose AI models can generate actionable knowledge on agroecological crop protection
- Academic journals' AI policies fail to curb the surge in AI-assisted academic writing
- Estimating the prevalence of LLM-assisted text in scholarly writing
- Writing in Symbiosis: Mapping Human Creative Agency in the AI Era
- AI-Assisted Writing Is Growing Fastest Among Non-English-Speaking and Less Established Scientists
- Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing
- Generative AI as a Linguistic Equalizer in Global Science
- Lexical Traces of <scp>AI</scp> : Linguistic Impact of Generative Tools in Academic Abstracts
- Return of the solo author: The changing division of labor in science in the age of generative AI
- The Rise of Large Language Models and the Direction and Impact of US Federal Research Funding
- A Survey of AI Scientists
- Visualizing the Adoption of Large Language Models across Sociology Subfields
- Artificial intelligence tools expand scientists’ impact but contract science’s focus
- Research Waste, Redundancy and the Rise of the Machines: The Questionable Future of Systematic Reviews
- From "Arbitrary Timberland" To "Skyline Charts": Is Visualization At Risk From The Pollution of Scientific Literature?
- The Potential for Artificial Intelligence to Augment Peer Review
- Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativity
- The diffusion of large language models in published academic articles
- Shifting norms in scholarly publications: trends in readability, objectivity, authorship, and AI use
- The Prompt Engineering Report Distilled: Quick Start Guide for Life Sciences
Discussions
- Delving into ChatGPT usage in academic writing through excess vocabulary [hn, 164 points, 105 comments]
- *Studying the "flowery" dialect of ChatGPT Delvish as it surreptitiously creeps into published scientific papers. #dialect #delvish #AIspeak arxiv.org/pdf/2406.07016 [bsky, 9 points, 1 comments]
- How nice to learn through this preprint that the verbosity and bombast of AI prose is already infecting scientific writing. (I like that "fluff" is coming into vogue as the descriptor for AI vocab.) . [bsky, 8 points, 1 comments]
- 📈 LLMs are quietly changing scientific writing. New research finds that 13.5% of biomedical abstracts in 2024 were likely processed with LLMs -- marked by spikes in the use of words like delves, cru [bsky, 7 points, 2 comments]
- Because if we, as academics, want "authentic assessments" for our students, then we'll need to acknowledge the fact that our own community uses AI for writing papers, or at least abstracts. Check out [bsky, 7 points, 0 comments]
- the uncanny valley-ness of AI text seems to go beyond vocabulary. there's also something repetitive & bland in the sentence structure & rhythm. It's one of those things I instinctively detect as "look [bsky, 7 points, 1 comments]
- Large Language Models have been used to edit 10% of abstracts in the last year!? "We study vocabulary changes in 14 million PubMed abstracts from 2010--2024, and show ... that at least 10% of 2024 abs [bsky, 5 points, 0 comments]
- "We study vocabulary changes in 14 million PubMed abstracts from 2010–2024, and show how the appearance of LLMs led to an abrupt increase in the frequency of certain style words. Our analysis...sugges [bsky, 4 points, 1 comments]
- Delving into ChatGPT usage in academic writing through excess vocabulary [hn, 2 points, 0 comments]
- As it turns out, "utilizing" and "leveraging" are both among a list of 291 "rare excess style words" which show a jump-like increase in frequency correlating with the emergence of LLMs as AI assistant [bsky, 2 points, 0 comments]
- Fascinating patterns have been discovered in how LLMs generate academic writing. Study here: arxiv.org/abs/2406.07016 [bsky, 2 points, 1 comments]
- @medericgc.bsky.social sur les mots véhiculés par les IA arxiv.org/abs/2406.07016 [bsky, 2 points, 1 comments]
- Here, www.science.org/doi/10.1126/..., the word I had in mind was "delve". Also, now, people are adapting to these "LLM" words and try to use others one with varying degree of success arxiv.org/html/2 [bsky, 2 points, 0 comments]
- Apparently about 10% of papers and maybe 30% in some sub-disciplines were written with assistance of LLMs, according to this study. arxiv.org/abs/2406.07016 That said, using LLMs to brush up prose is [bsky, 1 points, 2 comments]
- LLMs tend to over-use certain words, and I think that's part of what people pick up on. [bsky, 1 points, 0 comments]
- arxiv.org/abs/2406.07016 この論文っぽいけど、日経も何を見て使ったのか、簡単にでもいいからどういう手法で書かれた論文かくらい明記して欲しいな。 [bsky, 1 points, 0 comments]
- **Delving into LLM-assisted writing in biomedical publications through excess vocabulary** "_We study vocabulary changes in more than 15 million biomedical abstracts from 2010 to 2024 indexed by PubMe [mastodon, 0 points, 1 comments]
- ■「玉石混淆なネット上の集合知」を用いる危険性 例えば「ChatGPTの登場以降、なぜか英語論文において"delve"の単語が頻出するようになった」現象があり。 網羅的に調査されたプレプリント( arxiv.org/abs/2406.07016 )もある 例えば日本語入力するためのIME 無料で使えることから「Google日本語入力」を使ってる方も少なくないかと思うけど、あれはWeb上から学習し [bsky, 0 points, 1 comments]
- Intrigued by the extent to which generative AI is being used in the scientific literature? This fascinating paper, Delving into ChatGPT usage in academic writing, begins to offer answers. arxiv.org/pd [bsky, 0 points, 0 comments]
- An analysis of more than 15M PubMed abstracts (2010–2024) shows that the emergence of #LLM led to a sharp rise in the use of certain style words. At least 13.5% of abstracts published in 2024 were pro [bsky, 0 points, 0 comments]
- D'après cette étude - preprint - 10% des résumés / abstracts des publications scientifiques ont été générés avec #ChatGPT #journals [bsky, 0 points, 0 comments]
- Unter anderem wird es auch immer weniger lohnen, weil das Schreiben wissenschaftlicher Artikel selbst zunehmend automatisiert wird. Papermills on Steroids, sozusagen. Dazu gibt es schon eine Menge Ind [bsky, 0 points, 0 comments]
- these people know how to write ledes: > We show that the appearance of LLM-based writing assistants has had an unprecedented impact in the scientific literature, surpassing the effect of major world [bsky, 0 points, 0 comments]
- It's already everywhere: arxiv.org/abs/2406.07016 [bsky, 0 points, 0 comments]
Related