Can large language models provide useful feedback on research papers? A large-scale empirical analysis
2023/10/03 by Weixin Liang, Yuhui Zhang, Liang, Weixin +24 · 15 voices · 45 citations
Computer Science · Decision Sciences · Medicine · #Artificial Intelligence in Healthcare and Education #Scientific Computing and Data Management #Topic Modeling #cs.AI #cs.CL #cs.HC #cs.LG
paper · pdf · doi:10.48550/arxiv.2310.01783
openalex publication_date 2023/10/03 · openalex created_date 2023/10/05 · openalex updated_date 2026/07/28
Abstract
Expert feedback lays the foundation of rigorous research. However, the rapid growth of scholarly production and intricate knowledge specialization challenge the conventional scientific feedback mechanisms. High-quality peer reviews are increasingly difficult to obtain. Researchers who are more junior or from under-resourced settings have especially hard times getting timely feedback. With the breakthrough of large language models (LLM) such as GPT-4, there is growing interest in using LLMs to generate scientific feedback on research manuscripts. However, the utility of LLM-generated feedback has not been systematically studied. To address this gap, we created an automated pipeline using GPT-4 to provide comments on the full PDFs of scientific papers. We evaluated the quality of GPT-4's feedback through two large-scale studies. We first quantitatively compared GPT-4's generated feedback with human peer reviewer feedback in 15 Nature family journals (3,096 papers in total) and the ICLR machine learning conference (1,709 papers). The overlap in the points raised by GPT-4 and by human reviewers (average overlap 30.85% for Nature journals, 39.23% for ICLR) is comparable to the overlap between two human reviewers (average overlap 28.58% for Nature journals, 35.25% for ICLR). The overlap between GPT-4 and human reviewers is larger for the weaker papers. We then conducted a prospective user study with 308 researchers from 110 US institutions in the field of AI and computational biology to understand how researchers perceive feedback generated by our GPT-4 system on their own papers. Overall, more than half (57.4%) of the users found GPT-4 generated feedback helpful/very helpful and 82.4% found it more beneficial than feedback from at least some human reviewers. While our findings show that LLM-generated feedback can help researchers, we also identify several limitations.
Cited by
- Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers?
- From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines
- Does Multi-Agent Debate Improve AI Feedback on Research Papers?
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- AI Can Learn Scientific Taste
- Scientific production in the era of large language models
- Prestige over merit: An adapted audit of LLM bias in peer review
- Screening, sorting, and the feedback cycles that imperil peer review
- Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
- Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
- A Toolbox for Improving Evolutionary Prompt Search
- LLMs Can Assist with Proposal Selection at Large User Facilities
- Academic journals' AI policies fail to curb the surge in AI-assisted academic writing
- To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
- Can ChatGPT evaluate research environments? Evidence from REF2021
- FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
- OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
- A Cross‐Disciplinary Analysis of <scp>AI</scp> Policies in Academic Peer Review
- Beyond human gold standards: A multimodel framework for automated abstract classification and information extraction
- "Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers
- BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
- Product Manager Practices for Delegating Work to Generative AI: "Accountability must not be delegated to non-human actors"
- CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
- How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
- ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review
- AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
- Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review
- Paper Espresso: From Paper Overload to Research Insight
- Improving peer review capacity and integrity in forestry journals
- Análisis de los diálogos con ChatGPT del profesorado en formación durante el diseño de actividades didácticas de Química
- Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews
- Can GenAI Move from Individual Use to Collaborative Work? Experiences, Challenges, and Opportunities of Integrating GenAI into Collaborative Newsroom Routines
- Prompt Injection Attacks on LLM Generated Reviews of Scientific Publications
- When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
- The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
- Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
- MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
- A Multi-Task Evaluation of LLMs' Processing of Academic Text Input
- What are the limits to biomedical research acceleration through general-purpose AI?
- SlideAudit: A Dataset and Taxonomy for Automated Evaluation of Presentation Slides
- Oedipus and the Sphinx: Benchmarking and Improving Visual Language Models for Complex Graphic Reasoning
- Adaptive Cluster Collaborativeness Boosts LLMs Medical Decision Support Capacity
Discussions
- What could possibly go wrong ?
Large language models (LLM) to generate scientific feedback (peerreview) on research manuscripts
arxiv.org/abs/2310.01783 [bsky, 16 points, 4 comments]
- "The overlap in the points raised by GPT-4 and by human reviewers is comparable to the overlap between two human reviewers. 82.4% of the users found GPT-4 generated feedback more beneficial than feedb [bsky, 15 points, 2 comments]
- Editors - prepare for reviewers to just generate reviews using ChatGPT - that day is coming, soon.... arxiv.org/pdf/2310.017... [bsky, 9 points, 5 comments]
- "Can large language models provide useful feedback on research papers? A large-scale empirical analysis"
"Overlap in the points raised by GPT-4 and by human[s] ... is comparable to the overlap betwee [bsky, 6 points, 1 comments]
- Nearly 60% of #ChatGPT users think its feedback as a reviewer is useful, but over 80% think human feedback is still needed.
Main limitation of AI: lack of specificity. 1/
arxiv.org/abs/2310.01783 [bsky, 4 points, 2 comments]
- Can LLMs provide useful feedback on research papers? A broad empirical analysis [hn, 1 points, 0 comments]
- the source link OECD provided is arxiv.org/abs/2310.01783 not only is the data not there (in the form of a table or graph) this paper is also about LLM-generated feedback, not LLM-modified papers on t [bsky, 1 points, 0 comments]
- I just found out that they are: arxiv.org/pdf/2310.017...
"Overall, more than half (57.4%) of the users found GPT-4 generated feedback helpful/very helpful and 82.4% found it
more beneficial than feed [bsky, 1 points, 0 comments]
- 🧪 #AcademicSky #ml Direct link to the pre-print arxiv.org/abs/2310.01783 [bsky, 1 points, 0 comments]
- Some real research on this topic for contrast arxiv.org/pdf/2310.01783 / github.com/Weixin-Liang... [bsky, 1 points, 2 comments]
- LLMs as reviewers are as goog as Nature or ICLR -level peer reviewers! 😱
IMO this shows how poor is the peer review model we have today. 🤦♂️
arxiv.org/abs/2310.017... [bsky, 1 points, 0 comments]
- “We then conducted a prospective user study with 308 researchers from 110 US institutions…to understand how researchers perceive feedback generated by our GPT-4 system on their own papers” > “82.4% fo [bsky, 1 points, 1 comments]
- How well can #GPT4 provide scientific feedback on research projects? We study this in: arxiv.org/abs/2310.01783
We created a pipeline using GPT4 to read 1000s papers (from #Nature, #ICLR, etc.) and g [bsky, 0 points, 3 comments]
- And if you're not finding scholarly benefits to this technology it's because you're not looking or do not understand enough about the technology. arxiv.org/abs/2310.01783 [bsky, 0 points, 1 comments]
- Yet, finding qualified reviewers is a well-known challenge. On the other hand, the number of papers we write has been increasing at a drastic rate. "For example, the number of submissions to the ICLR [bsky, 0 points, 1 comments]
Related