Mathematical discoveries from program search with large language models
2023/12/14 by Bernardino Romera‐Paredes, Bernardino Romera-Paredes, Mohammadamin Barekatain +11 · 184 citations
Computer Science · #Computability, Logic, AI Algorithms #Artificial Intelligence in Games #Algorithms and Data Compression
paper · pdf · doi:10.1038/s41586-023-06924-6
Abstract
Abstract Large language models (LLMs) have demonstrated tremendous capabilities in solving complex tasks, from quantitative reasoning to understanding natural language. However, LLMs sometimes suffer from confabulations (or hallucinations), which can result in them making plausible but incorrect statements 1,2 . This hinders the use of current large models in scientific discovery. Here we introduce FunSearch (short for searching in the function space), an evolutionary procedure based on pairing a pretrained LLM with a systematic evaluator. We demonstrate the effectiveness of this approach to surpass the best-known results in important problems, pushing the boundary of existing LLM-based approaches 3 . Applying FunSearch to a central problem in extremal combinatorics—the cap set problem—we discover new constructions of large cap sets going beyond the best-known ones, both in finite dimensional and asymptotic cases. This shows that it is possible to make discoveries for established open problems using LLMs. We showcase the generality of FunSearch by applying it to an algorithmic problem, online bin packing, finding new heuristics that improve on widely used baselines. In contrast to most computer search approaches, FunSearch searches for programs that describe how to solve a problem, rather than what the solution is. Beyond being an effective and scalable strategy, discovered programs tend to be more interpretable than raw solutions, enabling feedback loops between domain experts and FunSearch, and the deployment of such programs in real-world applications.
Cited by
- CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
- Scattering Amplitudes as Programs: Self-Evolving Search for Theory and Event Generation
- Position: Quantum Program Generation Must Prioritize Validity Over Probabilistic Scaling
- FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
- The Café in Amsterdam: When the Incumbent Becomes the Oracle
- Language Model Crossover: Variation through Few-Shot Prompting
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- Symbolon: Symbolic Execution by Learning Code Transformation
- Autonomous Discovery of Wireless Communications Algorithms
- Program Synthesis for Simulation-Based Inference: Joint Model Selection and Parameter Estimation
- Automatic Construction of Clinical Scoring Systems with LLM Agents
- Automated Discovery Has No Universally Superior Harness
- Counting Cycles with AI: Counting Cycles with AI: Computationally Efficient Equivalent Forms with Applications
- From Heuristic Selection to Automated Algorithm Design: LLMs Benefit from Strong Priors
- Self-Modifying Lean Proof Agents with Verifier-Grounded Benchmark Coevolution
- B-repLer: Language-guided Editing of CAD Models
- Evolutionary Algorithm-Guided LLMs for Physics-Informed Neural Network Design
- Fantastic Adaptive Taxonomies and How to Use Them
- Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
- A Compositional Framework for Open-ended Intelligence
- OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining
- MILP-Evo: Closed-Loop Fully Automatic Design of MILP Solvers
- AI co-mathematician: Accelerating mathematicians with agentic AI
- Discovering Differences in Strategic Behavior Between Humans and LLMs
- Scientific production in the era of large language models
- AISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data Science
- FrontierCO: Real-World and Large-Scale Evaluation of Machine Learning Solvers for Combinatorial Optimization
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
- VLMaterial: Procedural Material Generation with Large Vision-Language Models
- Large Language Models Think Too Fast To Explore Effectively
- Reinforced Generation of Combinatorial Structures: Ramsey Numbers
- A Probabilistic Framework for LLM-Based Model Discovery
- LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs
- Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models
- Benchmarking Zero-Shot LLM-Generated Parent Selection in Genetic Programming for Symbolic Regression
- Bruhat intervals that are large hypercubes
- DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory
- Rethinking Logic Optimization Operators: Theory-Derived Operator Compression via Agentic Source Analysis
- MEMENTO: Memory-Guided Memetic Code-as-Policy Evolution
- A Vocabulary for Multi-Agent Automated Research Systems
- AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
- Tool-Augmented Hybrid Ensemble Reasoning with Distillation for Bilingual Mathematical Problem Solving
- Let the Barbarians In: How AI Can Accelerate Systems Performance Research
- PortAgent: LLM-driven Vehicle Dispatching Agent for Port Terminals
- Artificial Intelligence and Inherent Mathematical Difficulty
- EvoLattice: Persistent Internal-Population Evolution through Multi-Alternative Quality-Diversity Graph Representations for LLM-Guided Program Discovery
- Differentiable Evolutionary Reinforcement Learning
- Behavior and Representation in Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection
- Defining Cost Function of Steganography with Large Language Models
- CogMCTS: A Novel Cognitive-Guided Monte Carlo Tree Search Framework for Iterative Heuristic Evolution with Large Language Models
- AutoICE: Automatically Synthesizing Verifiable C Code via LLM-driven Evolution
- Model-Based and Sample-Efficient AI-Assisted Math Discovery in Sphere Packing
- auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation
- A Flexible Multi-Agent LLM-Human Framework for Fast Human Validated Tool Building
- ThetaEvolve: Test-time Learning on Open Problems
- Evolutionary Discovery of Heuristic Policies for Traffic Signal Control
- Automated Design Optimization via Strategic Search with Large Language Models
- Even with AI, Bijection Discovery is Still Hard: The Opportunities and Challenges of OpenEvolve for Novel Bijection Construction
- Cognitive Alpha Mining via LLM-Driven Code-Based Evolution
- MirrorMind: Empowering OmniScientist with the Expert Perspectives and Collective Knowledge of Human Scientists
- OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
- From Performance to Understanding: A Vision for Explainable Automated Algorithm Design
- Online Operator Design in Evolutionary Optimization for Flexible Job Shop Scheduling via Large Language Models
- GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms
- Cost-Driven Synthesis of Sound Abstract Interpreters
- Channel Ordering for Fairness in Elastic Optical Networks via a LLM-Guided Bottleneck TSP Solver
- irace-evo: Automatic Algorithm Configuration Extended With LLM-Based Code Evolution
- AgenticSciML: Collaborative Multi-Agent Systems for Emergent Discovery in Scientific Machine Learning
- Using Multi-modal Large Language Model to Boost Fireworks Algorithm's Ability in Settling Challenging Optimization Tasks
- Learning Interestingness in Automated Mathematical Theory Formation
- miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward
- Large Lemma Miners: Can LLMs do Induction Proofs for Hardware?
- Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network
- Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
- Reasoning Planning for Language Models
- ORGEval: Graph-Theoretic Evaluation of LLMs in Optimization Modeling
- SOCRATES: Simulation Optimization with Correlated Replicas and Adaptive Trajectory Evaluations
- PDE-SHARP: PDE Solver Hybrids through Analysis and Refinement Passes
- An In-depth Study of LLM Contributions to the Bin Packing Problem
- Glia: A Human-Inspired AI for Automated Systems Design and Optimization
- The FM Agent
- EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks
- Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?
- EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
- The social AI author: modeling creativity and distinction in simulated cultural fields
- Flows: Building Blocks of Reasoning and Collaborating AI
- optimizeanything: Unified Text Optimization can Outperform Specialized Systems
- STAR-PólyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision
- Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
- Understanding LoRA as Knowledge Memory: An Empirical Analysis
- Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
- Magellan: Autonomous Discovery of Novel Compiler Optimization Heuristics with AlphaEvolve
- FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data
- Discovering Heuristics with Large Language Models (LLMs) for Mixed-Integer Programs: Single-Machine Scheduling
- AI and the Decentering of Disciplinary Creativity
- Accelerating Materials Design via LLM-Guided Evolutionary Search
- REvolution: An Evolutionary Framework for RTL Generation driven by Large Language Models
- Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems
- KL-Regularized Reinforcement Learning is Designed to Mode Collapse
- An AI enhanced approach to the tree unimodality conjecture
- AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library
- EvoSyn: Generalizable Evolutionary Data Synthesis for Verifiable Learning
- Automated Algorithm Design for Auto-Tuning Optimizers
- An Agentic Framework with LLMs for Solving Complex Vehicle Routing Problems
- Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
- Programmatic Representation Learning with Language Models
- Where to Search: Measure the Prior-Structured Search Space of LLM Agents
- Thompson Sampling via Fine-Tuning of LLMs
- SR-Scientist: Scientific Equation Discovery With Agentic AI
- EvoCAD: Evolutionary CAD Code Generation with Vision Language Models
- Hierarchical Optimization via LLM-Guided Objective Evolution for Mobility-on-Demand Systems
- Mathematics with large language models as provers and verifiers
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- SIA: Self Improving AI with Harness & Weight Updates
- Meta-Harness: End-to-End Optimization of Model Harnesses
- Iterated Agent for Symbolic Regression
- VRPAgent: LLM-Driven Discovery of Heuristic Operators for Vehicle Routing Problems
- Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
- GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
- Scientific Algorithm Discovery by Augmenting AlphaEvolve with Deep Research
- MCCE: A Framework for Multi-LLM Collaborative Co-Evolution
- MetaMuse: Algorithm Generation via Creative Ideation
- EvoEngineer: Mastering Automated CUDA Kernel Code Evolution with Large Language Models
- LLM-Guided Evolutionary Program Synthesis for Quasi-Monte Carlo Design
- Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- EvoSpeak: Large Language Models for Interpretable Genetic Programming-Evolved Heuristics
- On Discovering Algorithms for Adversarial Imitation Learning
- Combining Large Language Models and Gradient-Free Optimization for Automatic Control Policy Synthesis
- Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
- Regression Language Models for Code
- Agentic Exploration of Physics Models
- Experience-Guided Reflective Co-Evolution of Prompts and Heuristics for Automatic Algorithm Design
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
- The impact of consistent internalization of the external effects of transport and manufacturing : a CGE analysis of Sweden
- TusoAI: Agentic Optimization for Scientific Methods
- FormalML: A Benchmark for Evaluating Formal Subgoal Completion in Machine Learning Theory
- Evaluating LLMs for Combinatorial Optimization: One-Phase and Two-Phase Heuristics for 2D Bin-Packing
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
- Bridging the Gap Between Scientific Laws Derived by AI Systems and Canonical Knowledge via Abductive Inference with AI-Noether
- GeoEvolve: Automating Geospatial Model Discovery via Multi-Agent Large Language Models
- Structuring Collective Action with LLM-Guided Evolution: From Ill-Structured Problems to Executable Heuristics
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
- Can AI Follow In Einstein's Footsteps?
- LLM-Guided Initialization for Accelerated Hybrid Quantum-Classical Medical Image Classification
- Optimizing ground state preparation protocols with autoresearch
- Lessons from complex systems science for AI governance
- CayleyPy Growth: Efficient growth computations and hundreds of new conjectures on Cayley graphs (Brief version)
- Reinforced Generation of Combinatorial Structures: Hardness of Approximation
- SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
- Large Language Models as End-to-end Combinatorial Optimization Solvers
- MetaGen: A DSL, Database, and Benchmark for VLM-Assisted Metamaterial Generation
- Improved Constructions and Lower Bounds for Maximally Recoverable Grid Codes
- Large Language Models in Operations Research: Methods, Applications, and Challenges
- Large Language Model Assisted Automated Algorithm Generation and Evolution via Meta-black-box optimization
- Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
- Evolution of Kernels: Automated RISC-V Kernel Optimization with Large Language Models
- EditDuet: A Multi-Agent System for Video Non-Linear Editing
- Autonomous Code Evolution Meets NP-Completeness
- Solve it with EASE
- LLM-Based Instance-Driven Heuristic Bias In the Context of a Biased Random Key Genetic Algorithm
- Artificial intelligence for representing and characterizing quantum systems
- Re-evaluating LLM-based Heuristic Search: A Case Study on the 3D Packing Problem
- Jointly Reinforcing Diversity and Quality in Language Model Generations
- Explicit Constructions of Maximal 3-Zero-Sum-Free Subsets in (ℤ/4ℤ)n
- Scaling Neuro-symbolic Problem Solving: Solver-Free Learning of Constraints and Objectives
- Computer-assisted graph theory: a survey
- Language Models For Generalised PDDL Planning: Synthesising Sound and Programmatic Policies
- ELATE: Evolutionary Language model for Automated Time-series Engineering
- VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization
- HiFo-Prompt: Prompting with Hindsight and Foresight for LLM-based Automatic Heuristic Design
- Discovering Expert-Level Nash Equilibrium Algorithms with Large Language Models
- EvoCut: Strengthening Integer Programs via Evolution-Guided Language Models
- Searching for Privacy Risks in LLM Agents via Simulation
- Route Planning and Online Routing for Quantum Key Distribution Networks
- MiGrATe: Mixed-Policy GRPO for Adaptation at Test-Time
- \(X\)-evolve: Solution space evolution powered by large language models
- The Missing Reward: Active Inference in the Era of Experience
- Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies
- Industrial LLM-based Code Optimization under Regulation: A Mixture-of-Agents Approach
- CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search
- How Far Are AI Scientists from Changing the World?
- Automatically discovering heuristics in a complex SAT solver with large language models
Related