Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
2024/09/06 by Chenglei Si, Diyi Yang, Si, Chenglei +3 · 23 voices · 77 citations
Social Sciences · #Wikis in Education and Collaboration #cs.AI #cs.CL #cs.CY #cs.HC #cs.LG
paper · pdf · doi:10.48550/arxiv.2409.04109
openalex publication_date 2024/09/06 · openalex created_date 2024/10/21 · openalex updated_date 2026/07/29
Abstract
Recent advancements in large language models (LLMs) have sparked optimism about their potential to accelerate scientific discovery, with a growing number of works proposing research agents that autonomously generate and validate new ideas. Despite this, no evaluations have shown that LLM systems can take the very first step of producing novel, expert-level ideas, let alone perform the entire research process. We address this by establishing an experimental design that evaluates research idea generation while controlling for confounders and performs the first head-to-head comparison between expert NLP researchers and an LLM ideation agent. By recruiting over 100 NLP researchers to write novel ideas and blind reviews of both LLM and human ideas, we obtain the first statistically significant conclusion on current LLM capabilities for research ideation: we find LLM-generated ideas are judged as more novel (p < 0.05) than human expert ideas while being judged slightly weaker on feasibility. Studying our agent baselines closely, we identify open problems in building and evaluating research agents, including failures of LLM self-evaluation and their lack of diversity in generation. Finally, we acknowledge that human judgements of novelty can be difficult, even by experts, and propose an end-to-end study design which recruits researchers to execute these ideas into full projects, enabling us to study whether these novelty and feasibility judgements result in meaningful differences in research outcome.
Cited by
Discussions
- Can LLMs Generate Novel Research Ideas? [hn, 50 points, 81 comments]
- Die Studie befasst sich mit der Frage: Können KI-Modelle originelle und neuartige Forschungsideen auf dem Niveau menschlicher Experten entwickeln? arxiv.org/abs/2409.04109 [bsky, 12 points, 2 comments]
- Inspired by the Can #LLMs Generate Novel #Research #Ideas #arxiv #paper - arxiv.org/abs/2409.04109 - I made a @poe.com app which tries to generate novel research ideas. Is it any good? That's beyond m [bsky, 6 points, 0 comments]
- Shot: LLMs are good at bullshit which is why they are increasingly being used to draft grants, which must hype or go bust Chaser: However, the projects incepted are hot air arxiv.org/abs/2409.04109 ar [bsky, 5 points, 1 comments]
- LLMs Outpace Humans in Novel Idea Generation [hn, 5 points, 0 comments]
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study [hn, 4 points, 0 comments]
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study [hn, 3 points, 0 comments]
- That hasn’t been my experience, we have multiple generations of open AI models whose cards detail a lack of progress in independent research and their self improvement project is a dead end. Where are [bsky, 2 points, 1 comments]
- But evidence shows different already. Here is a study showing that LLM generated ideas are judged as more novel than researcher generated ideas in NLP. arxiv.org/abs/2409.04109 [bsky, 2 points, 0 comments]
- Can LLMs Generate Novel Research Ideas? [hn, 2 points, 0 comments]
- ¿Pueden Claude o ChatGPT generar ideas de investigación novedosas? Un estudio de 2024 comparó propuestas de >100 investigadores en IA con las hechas por una IA mediante revisión ciega. El resultado: l [bsky, 1 points, 1 comments]
- I feel like the idea they're just input-output machines doesn't match how they actually work. They're not sentient or intelligent, but they're not simply databases. I've also seen a couple of studies [bsky, 1 points, 2 comments]
- Researchers have found that large language models #LLMs can generate research ideas deemed more novel than those from human experts (though slightly less feasible). #GenAI #Innovation #Research #Futur [bsky, 1 points, 1 comments]
- Can LLMs Generate Novel Research Ideas? New study reveals that large language models (LLMs) struggle to reliably evaluate ideas compared to human reviewers, with lower consistency in scores. 🤔 arxi [bsky, 1 points, 0 comments]
- Can LLMs Generate Novel Research Ideas? [bsky, 0 points, 0 comments]
- arxiv.org/abs/2409.04109 [bsky, 0 points, 0 comments]
- @avadeaux.bsky.social Din bild av LLMer verkar vara lite fördomsfull. https://arxiv.org/abs/2409.04109 [bsky, 0 points, 0 comments]
- Shot: LLMs are good at bullshit which is why they are increasingly being used to draft grants, which must hype or go bust Chaser: However, the projects incepted are hot air https://arxiv.org/abs/2409. [bsky, 0 points, 0 comments]
- LLMs now generate ML research proposals. Not sure whether they can do the same for social sciences, but perhaps more likely for quantitative research design? arxiv.org/abs/2409.041... [bsky, 0 points, 0 comments]
- Agents based on LLMs proposing machine learning research: "humans scored AI-generated and human-written proposals roughly equally in feasibility, expected effectiveness, how exciting they were, and ov [bsky, 0 points, 0 comments]
- Shot: LLMs are good at bullshit which is why they are increasingly being used to draft grants, which must hype or go bust Chaser: However, the projects incepted are hot air https://arxiv.org/abs/2409. [bsky, 0 points, 0 comments]
- Can LLMs Generate Novel Research Ideas? https://www.arxiv.org/abs/2409.04109 [bsky, 0 points, 0 comments]
- 1) In the narrow area of prompt generation techniques LLMs can generate ideas rated as more novel and exciting. They are sometimes less feasible. Out of 4000 ideas generated, only 200 were potentially [bsky, 0 points, 0 comments]
Related