Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
2024/09/06 by Chenglei Si, Diyi Yang, Si, Chenglei +3 · 23 voices · 160 citations
Computer Science · Social Sciences · #Artificial intelligence #Cartography #Computer science #Data science #Geography #Natural language processing #Scale (ratio) #Wikis in Education and Collaboration #cs.AI #cs.CL #cs.CY #cs.HC #cs.LG
paper · pdf · doi:10.48550/arxiv.2409.04109
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/09/06 · openalex created_date 2024/10/21 · openalex updated_date 2026/08/01
Abstract
Recent advancements in large language models (LLMs) have sparked optimism about their potential to accelerate scientific discovery, with a growing number of works proposing research agents that autonomously generate and validate new ideas. Despite this, no evaluations have shown that LLM systems can take the very first step of producing novel, expert-level ideas, let alone perform the entire research process. We address this by establishing an experimental design that evaluates research idea generation while controlling for confounders and performs the first head-to-head comparison between expert NLP researchers and an LLM ideation agent. By recruiting over 100 NLP researchers to write novel ideas and blind reviews of both LLM and human ideas, we obtain the first statistically significant conclusion on current LLM capabilities for research ideation: we find LLM-generated ideas are judged as more novel (p < 0.05) than human expert ideas while being judged slightly weaker on feasibility. Studying our agent baselines closely, we identify open problems in building and evaluating research agents, including failures of LLM self-evaluation and their lack of diversity in generation. Finally, we acknowledge that human judgements of novelty can be difficult, even by experts, and propose an end-to-end study design which recruits researchers to execute these ideas into full projects, enabling us to study whether these novelty and feasibility judgements result in meaningful differences in research outcome.
Cited by
- Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
- WARA: A Closed-Loop Multi-Agent Framework for Wireless Optimization Autoresearch
- ScienceMeter: Tracking Scientific Knowledge Updates in Language Models
- The unintended consequences of large language models as a labor-augmenting technology in science
- From Text to Discovery: How Large Language Models Are Reshaping Research Across Scientific and Humanistic Disciplines
- Before the Action: Benchmarking LLMs on Prospective Hypothesis Discovery
- Idea2Plan: Exploring AI-Powered Research Planning
- AI Research Agents Narrow Scientific Exploration
- Budgeted Subset Refinement for Execution-Aware LLM Research Ideation
- AI Can Learn Scientific Taste
- Job Anxiety in Post-Secondary Computer Science Students Caused by Artificial Intelligence
- Scientific production in the era of large language models
- Accumulating Context Changes the Beliefs of Language Models
- Epistemic Diversity and Knowledge Collapse in Large Language Models
- AISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data Science
- Paper2Code: Automating Code Generation from Scientific Papers in Machine Learning
- When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
- Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- Writing as a testbed for open ended agents
- All That Glitters is Not Novel: Plagiarism in AI Generated Research
- SWE-Lancer: Can Frontier LLMs Earn 1 Million from Real-World Freelance Software Engineering?
- Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
- AI as Entertainment
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- The Epistemological Consequences of Large Language Models: Rethinking collective intelligence and institutional knowledge
- TIB AIssistant: a Platform for AI-Supported Research Across Research Life Cycles
- The Erosion of LLM Signatures: Can We Still Distinguish Human and LLM-Generated Scientific Ideas After Iterative Paraphrasing?
- Mode-Conditioning Unlocks Superior Test-Time Scaling
- Assessing LLMs for Serendipity Discovery in Knowledge Graphs: A Case for Drug Repurposing
- Towards autonomous quantum physics research using LLM agents with access to intelligent tools
- On the Creativity of AI Agents
- Scientific judgment drifts over time in AI ideation
- Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
- Culture Cartography: Mapping the Landscape of Cultural Knowledge
- Diffuse Thinking: Exploring Diffusion Language Models as Efficient Thought Proposers for Reasoning
- One Run Is Not an Idea: The Implementation Lottery in Automated Research
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Towards AI as Colleagues: Multi-Agent System Improves Structured Professional Ideation
- How Do AI Agents Do Human Work? Comparing AI and Human Workflows Across Diverse Occupations
- Magellan: Guided MCTS for Latent Space Exploration and Novelty Generation
- A computational model and tool for generating more novel opportunities in professional innovation processes
- Black Box Absorption: LLMs Undermining Innovative Ideas
- TrustResearcher: Automating Knowledge-Grounded and Transparent Research Ideation with Multi-Agent Collaboration
- ScholarEval: Research Idea Evaluation Grounded in Literature
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models
- FML-bench: A Benchmark for Automatic ML Research Agents Highlighting the Importance of Exploration Breadth
- An Alternative Trajectory for Generative AI
- BILLY: Steering Large Language Models via Merging Persona Vectors for Creative Generation
- IoDResearch: Deep Research on Private Heterogeneous Data via the Internet of Data
- Evolving and Executing Research Plans via Double-Loop Multi-Agent Collaboration
- TinyScientist: An Interactive, Extensible, and Controllable Framework for Building Research Agents
- RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback
- Scientific Algorithm Discovery by Augmenting AlphaEvolve with Deep Research
- Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
- NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality Estimation
- Beyond efficiency: How artificial intelligence (AI) will reshape scientific inquiry and the publication process
- Exploring the scope of generative AI in literature review development
- MotivGraph-SoIQ: Integrating Motivational Knowledge Graphs and Socratic Dialogue for Enhanced LLM Ideation
- FlexMind: Supporting Deeper Creative Thinking with LLMs
- TCA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval
- Diversifying Personalized Research Ideation against AI-Induced Homogenization
- Can AI Follow In Einstein's Footsteps?
- Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology
- OpenLens AI: Fully Autonomous Research Agent for Health Infomatics
- Scholarly Communications in 2025: An Aerial Evaluation of a System Challenged by AI and Much More
- Jointly Reinforcing Diversity and Quality in Language Model Generations
- ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
- What Are Research Hypotheses?
- The Ramon Llull's Thinking Machine for Automated Ideation
- Effective Red-Teaming of Policy-Adherent Agents
- AINL-Eval 2025 Shared Task: Detection of AI-Generated Scientific Abstracts in Russian
- Synthesizing scientific literature with retrieval-augmented language models
- ConlangCrafter: Constructing Languages with a Multi-Hop LLM Pipeline
- Beyond Brainstorming: What Drives High-Quality Scientific Ideas? Lessons from Multi-Agent Collaboration
- MaRGen: Multi-Agent LLM Approach for Self-Directed Market Research and Analysis
- How Far Are AI Scientists from Changing the World?
- IDRBench: Understanding the Capability of Large Language Models on Interdisciplinary Research
- Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics
- RobotValues: Evaluating Household Robots When Human Values Conflict
- IRPAPERS: A Visual Document Benchmark for Scientific Retrieval and Question Answering
- Towards Execution-Grounded Automated AI Research
- Combinatorial Creativity: A New Frontier in Generalization Abilities
- AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research
- Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and Evaluator
- Exploring Design of Multi-Agent LLM Dialogues for Research Ideation
- MK2 at PBIG Competition: A Prompt Generation Solution
- Abductive Computational Systems: Creative Abduction and Future Directions
- AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research
- Scaffolding Recursive Divergence and Convergence in Story Ideation
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
- Agent Ideate: A Framework for Product Idea Generation from Patents Using Agentic AI
- SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
- RExBench: Can coding agents autonomously implement AI research extensions?
- Literature-Grounded Novelty Assessment of Scientific Ideas
- THE-Tree: Can Tracing Historical Evolution Enhance Scientific Verification and Reasoning?
- Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
- LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- HypER: Literature-grounded Hypothesis Generation and Distillation with Provenance
- Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback
- Formalizing Learning from Language Feedback with Provable Guarantees
- From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer Review
- Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
- When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
- Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Science
- Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework
- AI Scientists Fail Without Strong Implementation Capability
- ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code
- Predicting Empirical AI Research Outcomes with Language Models
- MIR: Methodology Inspiration Retrieval for Scientific Research Problems
- Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey
- From Reasoning to Learning: A Survey on Hypothesis Discovery and Rule Learning with Large Language Models
- Augmenting Research Ideation with Data: An Empirical Investigation in Social Science
- CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation
- AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
- AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy
- MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
- MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search
- OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
- AI-Researcher: Autonomous Scientific Innovation
- Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
- MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback
- Generative AI and Creativity: A Systematic Literature Review and Meta-Analysis
- Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation
- Long-Horizon Autonomous Architecture Research with a Language-Model Agent: A Behavioural Case Study
- Toward Reliable Scientific Hypothesis Generation: Evaluating Truthfulness and Hallucination in Large Language Models
- WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications
- Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models
- From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
- XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration
- Assessing the Effect of Cross-Domain Mapping on Creativity in Humans and Large Language Models
- ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment
- Vibe Researching as Wolf Coming: Can AI Agents with Skills Replace or Augment Social Scientists?
- OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Algorithm Discovery
- Measuring the Gap Between Human and LLM Research Ideas
- PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement
- SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
- Why LLMs Aren't Scientists Yet: Lessons from Four Autonomous Research Attempts
- Towards Automated Scoping of AI for Social Good Projects
- Teaching Language Models to Forecast Research Success Through Comparative Idea Evaluation
- Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities
- Transformational Creativity in Science: A Graphical Theory
- Spark: A System for Scientifically Creative Idea Generation
- AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions
- Ensemble Bayesian Inference: Leveraging Small Language Models to Achieve LLM-level Accuracy in Profile Matching Tasks
- IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery
- ArXivBench: When You Should Avoid Using ChatGPT for Academic Writing
- Aspirational Affordances of AI
- Sparks of Science: Hypothesis Generation Using Structured Paper Data
- AI Safety Should Prioritize the Future of Work
- Resurrecting Socrates in the Age of AI: A Study Protocol for Evaluating a Socratic Tutor to Support Research Question Development in Higher Education
- An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics
- Reimagining Urban Science: Scaling Causal Inference with Large Language Models
- HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
- Has the Creativity of Large-Language Models peaked? An analysis of inter- and intra-LLM variability
- Societal Impacts Research Requires Benchmarks for Creative Composition Tasks
Discussions
- Can LLMs Generate Novel Research Ideas? [hn, 50 points, 81 comments]
- Die Studie befasst sich mit der Frage: Können KI-Modelle originelle und neuartige Forschungsideen auf dem Niveau menschlicher Experten entwickeln? arxiv.org/abs/2409.04109 [bsky, 12 points, 2 comments]
- Inspired by the Can #LLMs Generate Novel #Research #Ideas #arxiv #paper - arxiv.org/abs/2409.04109 - I made a @poe.com app which tries to generate novel research ideas. Is it any good? That's beyond m [bsky, 6 points, 0 comments]
- Shot: LLMs are good at bullshit which is why they are increasingly being used to draft grants, which must hype or go bust Chaser: However, the projects incepted are hot air arxiv.org/abs/2409.04109 ar [bsky, 5 points, 1 comments]
- LLMs Outpace Humans in Novel Idea Generation [hn, 5 points, 0 comments]
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study [hn, 4 points, 0 comments]
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study [hn, 3 points, 0 comments]
- That hasn’t been my experience, we have multiple generations of open AI models whose cards detail a lack of progress in independent research and their self improvement project is a dead end. Where are [bsky, 2 points, 1 comments]
- But evidence shows different already. Here is a study showing that LLM generated ideas are judged as more novel than researcher generated ideas in NLP. arxiv.org/abs/2409.04109 [bsky, 2 points, 0 comments]
- Can LLMs Generate Novel Research Ideas? [hn, 2 points, 0 comments]
- ¿Pueden Claude o ChatGPT generar ideas de investigación novedosas? Un estudio de 2024 comparó propuestas de >100 investigadores en IA con las hechas por una IA mediante revisión ciega. El resultado: l [bsky, 1 points, 1 comments]
- I feel like the idea they're just input-output machines doesn't match how they actually work. They're not sentient or intelligent, but they're not simply databases. I've also seen a couple of studies [bsky, 1 points, 2 comments]
- Researchers have found that large language models #LLMs can generate research ideas deemed more novel than those from human experts (though slightly less feasible). #GenAI #Innovation #Research #Futur [bsky, 1 points, 1 comments]
- Can LLMs Generate Novel Research Ideas? New study reveals that large language models (LLMs) struggle to reliably evaluate ideas compared to human reviewers, with lower consistency in scores. 🤔 arxi [bsky, 1 points, 0 comments]
- Can LLMs Generate Novel Research Ideas? [bsky, 0 points, 0 comments]
- arxiv.org/abs/2409.04109 [bsky, 0 points, 0 comments]
- @avadeaux.bsky.social Din bild av LLMer verkar vara lite fördomsfull. https://arxiv.org/abs/2409.04109 [bsky, 0 points, 0 comments]
- Shot: LLMs are good at bullshit which is why they are increasingly being used to draft grants, which must hype or go bust Chaser: However, the projects incepted are hot air https://arxiv.org/abs/2409. [bsky, 0 points, 0 comments]
- LLMs now generate ML research proposals. Not sure whether they can do the same for social sciences, but perhaps more likely for quantitative research design? arxiv.org/abs/2409.041... [bsky, 0 points, 0 comments]
- Agents based on LLMs proposing machine learning research: "humans scored AI-generated and human-written proposals roughly equally in feasibility, expected effectiveness, how exciting they were, and ov [bsky, 0 points, 0 comments]
- Shot: LLMs are good at bullshit which is why they are increasingly being used to draft grants, which must hype or go bust Chaser: However, the projects incepted are hot air https://arxiv.org/abs/2409. [bsky, 0 points, 0 comments]
- Can LLMs Generate Novel Research Ideas? https://www.arxiv.org/abs/2409.04109 [bsky, 0 points, 0 comments]
- 1) In the narrow area of prompt generation techniques LLMs can generate ideas rated as more novel and exciting. They are sometimes less feasible. Out of 4000 ideas generated, only 200 were potentially [bsky, 0 points, 0 comments]
Related