Can large language models provide useful feedback on research papers? A large-scale empirical analysis
2023/10/03 by Weixin Liang, Yuhui Zhang, Liang, Weixin +24 · 15 voices · 87 citations
Computer Science · Decision Sciences · Engineering · Mathematics · Medicine · Psychology · #Artificial Intelligence in Healthcare and Education #Computer science #Data science #Empirical research #Engineering #Feedback control #Field (mathematics) #Geography #Mathematics #Mathematics education #Peer feedback #Peer review #Pipeline (software) #Political science #Psychology #Quality (philosophy) #Scale (ratio) #Scientific Computing and Data Management #Statistics #Topic Modeling #cs.AI #cs.CL #cs.HC #cs.LG
paper · pdf · doi:10.48550/arxiv.2310.01783
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/10/03 · openalex created_date 2023/10/05 · openalex updated_date 2026/08/06
Abstract
Expert feedback lays the foundation of rigorous research. However, the rapid growth of scholarly production and intricate knowledge specialization challenge the conventional scientific feedback mechanisms. High-quality peer reviews are increasingly difficult to obtain. Researchers who are more junior or from under-resourced settings have especially hard times getting timely feedback. With the breakthrough of large language models (LLM) such as GPT-4, there is growing interest in using LLMs to generate scientific feedback on research manuscripts. However, the utility of LLM-generated feedback has not been systematically studied. To address this gap, we created an automated pipeline using GPT-4 to provide comments on the full PDFs of scientific papers. We evaluated the quality of GPT-4's feedback through two large-scale studies. We first quantitatively compared GPT-4's generated feedback with human peer reviewer feedback in 15 Nature family journals (3,096 papers in total) and the ICLR machine learning conference (1,709 papers). The overlap in the points raised by GPT-4 and by human reviewers (average overlap 30.85% for Nature journals, 39.23% for ICLR) is comparable to the overlap between two human reviewers (average overlap 28.58% for Nature journals, 35.25% for ICLR). The overlap between GPT-4 and human reviewers is larger for the weaker papers. We then conducted a prospective user study with 308 researchers from 110 US institutions in the field of AI and computational biology to understand how researchers perceive feedback generated by our GPT-4 system on their own papers. Overall, more than half (57.4%) of the users found GPT-4 generated feedback helpful/very helpful and 82.4% found it more beneficial than feedback from at least some human reviewers. While our findings show that LLM-generated feedback can help researchers, we also identify several limitations.
Cited by
- Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers?
- From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines
- Does Multi-Agent Debate Improve AI Feedback on Research Papers?
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- AI Can Learn Scientific Taste
- Scientific production in the era of large language models
- Prestige over merit: An adapted audit of LLM bias in peer review
- Screening, sorting, and the feedback cycles that imperil peer review
- Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
- Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
- A Toolbox for Improving Evolutionary Prompt Search
- LLMs Can Assist with Proposal Selection at Large User Facilities
- Academic journals' AI policies fail to curb the surge in AI-assisted academic writing
- To Err Is Human: Systematic Quantification of Errors in Published AI Papers via LLM Analysis
- Can ChatGPT evaluate research environments? Evidence from REF2021
- FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
- OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
- A Cross‐Disciplinary Analysis of AI Policies in Academic Peer Review
- Beyond human gold standards: A multimodel framework for automated abstract classification and information extraction
- "Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers
- BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
- Product Manager Practices for Delegating Work to Generative AI: "Accountability must not be delegated to non-human actors"
- CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
- How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
- ReviewerToo: Should AI Join The Program Committee? A Look At The Future of Peer Review
- AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
- Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review
- Paper Espresso: From Paper Overload to Research Insight
- Improving peer review capacity and integrity in forestry journals
- Análisis de los diálogos con ChatGPT del profesorado en formación durante el diseño de actividades didácticas de Química
- Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews
- Can GenAI Move from Individual Use to Collaborative Work? Experiences, Challenges, and Opportunities of Integrating GenAI into Collaborative Newsroom Routines
- Prompt Injection Attacks on LLM Generated Reviews of Scientific Publications
- When Your Reviewer is an LLM: Biases, Divergence, and Prompt Injection Risks in Peer Review
- The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
- Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
- MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
- A Multi-Task Evaluation of LLMs' Processing of Academic Text Input
- What are the limits to biomedical research acceleration through general-purpose AI?
- SlideAudit: A Dataset and Taxonomy for Automated Evaluation of Presentation Slides
- Oedipus and the Sphinx: Benchmarking and Improving Visual Language Models for Complex Graphic Reasoning
- Adaptive Cluster Collaborativeness Boosts LLMs Medical Decision Support Capacity
- How do authors want to use AI for review?
- Towards Execution-Grounded Automated AI Research
- Peer Review and the Diffusion of Ideas
- Psychology-Driven Enhancement of Humour Translation
- A Survey of Pun Generation: Datasets, Evaluations and Methodologies
- The Trends of Open Access Academic Books and Discipline Dynamics: A Cross-database Comparison Based on OpenAlex and Web of Science
- Position: The ML Community Must Build an AI-Augmented Peer-Review Ecosystem
- TreeReview: A Dynamic Tree of Questions Framework for Deep and Efficient LLM-based Scientific Peer Review
- Can Large Language Models Be Trusted Paper Reviewers? A Feasibility Study
- Breaking the Reviewer: Assessing the Vulnerability of Large Language Models in Automated Peer Review Under Textual Adversarial Attacks
- From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer Review
- Exploring the Potential of Large Language Models in Differential Abundance Analysis
- Evaluating Large Language Model Capabilities in Assessing Spatial Econometrics Research
- ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
- Predicting Empirical AI Research Outcomes with Language Models
- Reviewing Scientific Papers for Critical Problems With Reasoning LLMs: Baseline Approaches and Automatic Evaluation
- Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development
- Text2Grad: Reinforcement Learning from Natural Language Feedback
- Language Models Should be Used to Surface the Unwritten Code of Science and Society
- OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
- Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery
- SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers
- Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
- XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration
- REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning
- ReVoicer: Conversational Voice Annotation for Human-Centered, LLM-Assisted Peer Review
- RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Review
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- FOXGLOVE: Understanding Goal-Oriented and Anchored Writing Feedback from Experts and LLMs on Argumentative Essays
- Towards a new paradigm of scientific discovery with socialized artificial intelligence
- ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration
- GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
- NovBench: Evaluating Large Language Models on Academic Paper Novelty Assessment
- Large Language Models for Departmental Expert Review Quality Scores
- Reward Modeling for Scientific Writing Evaluation
- AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality
- In which fields do ChatGPT 4o scores align better than citations with research quality?
- A Vision for the Future of an AI-Integrated Research Ecosystem
- Reimagining Urban Science: Scaling Causal Inference with Large Language Models
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images
- Identifying Aspects in Peer Reviews
Discussions
- What could possibly go wrong ?
Large language models (LLM) to generate scientific feedback (peerreview) on research manuscripts
arxiv.org/abs/2310.01783 [bsky, 16 points, 4 comments]
- "The overlap in the points raised by GPT-4 and by human reviewers is comparable to the overlap between two human reviewers. 82.4% of the users found GPT-4 generated feedback more beneficial than feedb [bsky, 15 points, 2 comments]
- Editors - prepare for reviewers to just generate reviews using ChatGPT - that day is coming, soon.... arxiv.org/pdf/2310.017... [bsky, 9 points, 5 comments]
- "Can large language models provide useful feedback on research papers? A large-scale empirical analysis"
"Overlap in the points raised by GPT-4 and by human[s] ... is comparable to the overlap betwee [bsky, 6 points, 1 comments]
- Nearly 60% of #ChatGPT users think its feedback as a reviewer is useful, but over 80% think human feedback is still needed.
Main limitation of AI: lack of specificity. 1/
arxiv.org/abs/2310.01783 [bsky, 4 points, 2 comments]
- Can LLMs provide useful feedback on research papers? A broad empirical analysis [hn, 1 points, 0 comments]
- the source link OECD provided is arxiv.org/abs/2310.01783 not only is the data not there (in the form of a table or graph) this paper is also about LLM-generated feedback, not LLM-modified papers on t [bsky, 1 points, 0 comments]
- I just found out that they are: arxiv.org/pdf/2310.017...
"Overall, more than half (57.4%) of the users found GPT-4 generated feedback helpful/very helpful and 82.4% found it
more beneficial than feed [bsky, 1 points, 0 comments]
- 🧪 #AcademicSky #ml Direct link to the pre-print arxiv.org/abs/2310.01783 [bsky, 1 points, 0 comments]
- Some real research on this topic for contrast arxiv.org/pdf/2310.01783 / github.com/Weixin-Liang... [bsky, 1 points, 2 comments]
- LLMs as reviewers are as goog as Nature or ICLR -level peer reviewers! 😱
IMO this shows how poor is the peer review model we have today. 🤦♂️
arxiv.org/abs/2310.017... [bsky, 1 points, 0 comments]
- “We then conducted a prospective user study with 308 researchers from 110 US institutions…to understand how researchers perceive feedback generated by our GPT-4 system on their own papers” > “82.4% fo [bsky, 1 points, 1 comments]
- How well can #GPT4 provide scientific feedback on research projects? We study this in: arxiv.org/abs/2310.01783
We created a pipeline using GPT4 to read 1000s papers (from #Nature, #ICLR, etc.) and g [bsky, 0 points, 3 comments]
- And if you're not finding scholarly benefits to this technology it's because you're not looking or do not understand enough about the technology. arxiv.org/abs/2310.01783 [bsky, 0 points, 1 comments]
- Yet, finding qualified reviewers is a well-known challenge. On the other hand, the number of papers we write has been increasing at a drastic rate. "For example, the number of submissions to the ICLR [bsky, 0 points, 1 comments]
Related