Mathematical discoveries from program search with large language models
2023/12/14 by Bernardino Romera‐Paredes, Bernardino Romera-Paredes, Mohammadamin Barekatain +11 · 390 citations
Computer Science · #Algorithms and Data Compression #Artificial Intelligence in Games #Computability, Logic, AI Algorithms #Computer science #Data science #Programming language
paper · pdf · doi:10.1038/s41586-023-06924-6
published in Nature 625(7995), 468-475 (Nature Portfolio)
openalex publication_date 2023/12/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
Abstract
Abstract Large language models (LLMs) have demonstrated tremendous capabilities in solving complex tasks, from quantitative reasoning to understanding natural language. However, LLMs sometimes suffer from confabulations (or hallucinations), which can result in them making plausible but incorrect statements 1,2 . This hinders the use of current large models in scientific discovery. Here we introduce FunSearch (short for searching in the function space), an evolutionary procedure based on pairing a pretrained LLM with a systematic evaluator. We demonstrate the effectiveness of this approach to surpass the best-known results in important problems, pushing the boundary of existing LLM-based approaches 3 . Applying FunSearch to a central problem in extremal combinatorics—the cap set problem—we discover new constructions of large cap sets going beyond the best-known ones, both in finite dimensional and asymptotic cases. This shows that it is possible to make discoveries for established open problems using LLMs. We showcase the generality of FunSearch by applying it to an algorithmic problem, online bin packing, finding new heuristics that improve on widely used baselines. In contrast to most computer search approaches, FunSearch searches for programs that describe how to solve a problem, rather than what the solution is. Beyond being an effective and scalable strategy, discovered programs tend to be more interpretable than raw solutions, enabling feedback loops between domain experts and FunSearch, and the deployment of such programs in real-world applications.
Cited by
- CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
- Scattering Amplitudes as Programs: Self-Evolving Search for Theory and Event Generation
- Position: Quantum Program Generation Must Prioritize Validity Over Probabilistic Scaling
- FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
- The Café in Amsterdam: When the Incumbent Becomes the Oracle
- Language Model Crossover: Variation through Few-Shot Prompting
- Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
- Symbolon: Symbolic Execution by Learning Code Transformation
- Autonomous Discovery of Wireless Communications Algorithms
- Program Synthesis for Simulation-Based Inference: Joint Model Selection and Parameter Estimation
- Automatic Construction of Clinical Scoring Systems with LLM Agents
- Automated Discovery Has No Universally Superior Harness
- Counting Cycles with AI: Counting Cycles with AI: Computationally Efficient Equivalent Forms with Applications
- From Heuristic Selection to Automated Algorithm Design: LLMs Benefit from Strong Priors
- Self-Modifying Lean Proof Agents with Verifier-Grounded Benchmark Coevolution
- B-repLer: Language-guided Editing of CAD Models
- Evolutionary Algorithm-Guided LLMs for Physics-Informed Neural Network Design
- Fantastic Adaptive Taxonomies and How to Use Them
- Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
- A Compositional Framework for Open-ended Intelligence
- OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining
- MILP-Evo: Closed-Loop Fully Automatic Design of MILP Solvers
- AI co-mathematician: Accelerating mathematicians with agentic AI
- Discovering Differences in Strategic Behavior Between Humans and LLMs
- Scientific production in the era of large language models
- AISSISTANT: Human-AI Collaborative Review and Perspective Research Workflows in Data Science
- FrontierCO: Real-World and Large-Scale Evaluation of Machine Learning Solvers for Combinatorial Optimization
- Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
- How Deep Do Large Language Models Internalize Scientific Literature and Citation Practices?
- Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?
- VLMaterial: Procedural Material Generation with Large Vision-Language Models
- Large Language Models Think Too Fast To Explore Effectively
- Reinforced Generation of Combinatorial Structures: Ramsey Numbers
- A Probabilistic Framework for LLM-Based Model Discovery
- LLM4Branch: Large Language Model for Discovering Efficient Branching Policies of Integer Programs
- Code-Space Response Oracles: Generating Interpretable Multi-Agent Policies with Large Language Models
- Benchmarking Zero-Shot LLM-Generated Parent Selection in Genetic Programming for Symbolic Regression
- Bruhat intervals that are large hypercubes
- DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory
- Rethinking Logic Optimization Operators: Theory-Derived Operator Compression via Agentic Source Analysis
- MEMENTO: Memory-Guided Memetic Code-as-Policy Evolution
- A Vocabulary for Multi-Agent Automated Research Systems
- AInsteinBench: Benchmarking Coding Agents on Scientific Repositories
- Tool-Augmented Hybrid Ensemble Reasoning with Distillation for Bilingual Mathematical Problem Solving
- Let the Barbarians In: How AI Can Accelerate Systems Performance Research
- PortAgent: LLM-driven Vehicle Dispatching Agent for Port Terminals
- Artificial Intelligence and Inherent Mathematical Difficulty
- EvoLattice: Persistent Internal-Population Evolution through Multi-Alternative Quality-Diversity Graph Representations for LLM-Guided Program Discovery
- Differentiable Evolutionary Reinforcement Learning
- Behavior and Representation in Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection
- Defining Cost Function of Steganography with Large Language Models
- CogMCTS: A Novel Cognitive-Guided Monte Carlo Tree Search Framework for Iterative Heuristic Evolution with Large Language Models
- AutoICE: Automatically Synthesizing Verifiable C Code via LLM-driven Evolution
- Model-Based and Sample-Efficient AI-Assisted Math Discovery in Sphere Packing
- auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation
- A Flexible Multi-Agent LLM-Human Framework for Fast Human Validated Tool Building
- ThetaEvolve: Test-time Learning on Open Problems
- Evolutionary Discovery of Heuristic Policies for Traffic Signal Control
- Automated Design Optimization via Strategic Search with Large Language Models
- Even with AI, Bijection Discovery is Still Hard: The Opportunities and Challenges of OpenEvolve for Novel Bijection Construction
- Cognitive Alpha Mining via LLM-Driven Code-Based Evolution
- MirrorMind: Empowering OmniScientist with the Expert Perspectives and Collective Knowledge of Human Scientists
- OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
- From Performance to Understanding: A Vision for Explainable Automated Algorithm Design
- Online Operator Design in Evolutionary Optimization for Flexible Job Shop Scheduling via Large Language Models
- GigaEvo: An Open Source Optimization Framework Powered By LLMs And Evolution Algorithms
- Cost-Driven Synthesis of Sound Abstract Interpreters
- Channel Ordering for Fairness in Elastic Optical Networks via a LLM-Guided Bottleneck TSP Solver
- irace-evo: Automatic Algorithm Configuration Extended With LLM-Based Code Evolution
- AgenticSciML: Collaborative Multi-Agent Systems for Emergent Discovery in Scientific Machine Learning
- Using Multi-modal Large Language Model to Boost Fireworks Algorithm's Ability in Settling Challenging Optimization Tasks
- Learning Interestingness in Automated Mathematical Theory Formation
- miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward
- Large Lemma Miners: Can LLMs do Induction Proofs for Hardware?
- Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network
- Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
- Reasoning Planning for Language Models
- ORGEval: Graph-Theoretic Evaluation of LLMs in Optimization Modeling
- SOCRATES: Simulation Optimization with Correlated Replicas and Adaptive Trajectory Evaluations
- PDE-SHARP: PDE Solver Hybrids through Analysis and Refinement Passes
- An In-depth Study of LLM Contributions to the Bin Packing Problem
- Glia: A Human-Inspired AI for Automated Systems Design and Optimization
- The FM Agent
- EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks
- Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?
- EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
- The social AI author: modeling creativity and distinction in simulated cultural fields
- Flows: Building Blocks of Reasoning and Collaborating AI
- optimizeanything: Unified Text Optimization can Outperform Specialized Systems
- STAR-PólyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision
- Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
- Understanding LoRA as Knowledge Memory: An Empirical Analysis
- Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
- Magellan: Autonomous Discovery of Novel Compiler Optimization Heuristics with AlphaEvolve
- FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data
- Discovering Heuristics with Large Language Models (LLMs) for Mixed-Integer Programs: Single-Machine Scheduling
- AI and the Decentering of Disciplinary Creativity
- Accelerating Materials Design via LLM-Guided Evolutionary Search
- REvolution: An Evolutionary Framework for RTL Generation driven by Large Language Models
- Co-Designing Quantum Codes with Transversal Diagonal Gates via Multi-Agent Systems
- KL-Regularized Reinforcement Learning is Designed to Mode Collapse
- An AI enhanced approach to the tree unimodality conjecture
- AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library
- EvoSyn: Generalizable Evolutionary Data Synthesis for Verifiable Learning
- Automated Algorithm Design for Auto-Tuning Optimizers
- An Agentic Framework with LLMs for Solving Complex Vehicle Routing Problems
- Foundation Models for Scientific Discovery: From Paradigm Enhancement to Paradigm Transition
- Programmatic Representation Learning with Language Models
- Where to Search: Measure the Prior-Structured Search Space of LLM Agents
- Thompson Sampling via Fine-Tuning of LLMs
- SR-Scientist: Scientific Equation Discovery With Agentic AI
- EvoCAD: Evolutionary CAD Code Generation with Vision Language Models
- Hierarchical Optimization via LLM-Guided Objective Evolution for Mobility-on-Demand Systems
- Mathematics with large language models as provers and verifiers
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- SIA: Self Improving AI with Harness & Weight Updates
- Meta-Harness: End-to-End Optimization of Model Harnesses
- Iterated Agent for Symbolic Regression
- VRPAgent: LLM-Driven Discovery of Heuristic Operators for Vehicle Routing Problems
- Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
- GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
- Scientific Algorithm Discovery by Augmenting AlphaEvolve with Deep Research
- MCCE: A Framework for Multi-LLM Collaborative Co-Evolution
- MetaMuse: Algorithm Generation via Creative Ideation
- EvoEngineer: Mastering Automated CUDA Kernel Code Evolution with Large Language Models
- LLM-Guided Evolutionary Program Synthesis for Quasi-Monte Carlo Design
- Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- EvoSpeak: Large Language Models for Interpretable Genetic Programming-Evolved Heuristics
- On Discovering Algorithms for Adversarial Imitation Learning
- Combining Large Language Models and Gradient-Free Optimization for Automatic Control Policy Synthesis
- Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
- Regression Language Models for Code
- Agentic Exploration of Physics Models
- Experience-Guided Reflective Co-Evolution of Prompts and Heuristics for Automatic Algorithm Design
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
- The impact of consistent internalization of the external effects of transport and manufacturing : a CGE analysis of Sweden
- TusoAI: Agentic Optimization for Scientific Methods
- FormalML: A Benchmark for Evaluating Formal Subgoal Completion in Machine Learning Theory
- Evaluating LLMs for Combinatorial Optimization: One-Phase and Two-Phase Heuristics for 2D Bin-Packing
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
- Bridging the Gap Between Scientific Laws Derived by AI Systems and Canonical Knowledge via Abductive Inference with AI-Noether
- GeoEvolve: Automating Geospatial Model Discovery via Multi-Agent Large Language Models
- Structuring Collective Action with LLM-Guided Evolution: From Ill-Structured Problems to Executable Heuristics
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
- Can AI Follow In Einstein's Footsteps?
- LLM-Guided Initialization for Accelerated Hybrid Quantum-Classical Medical Image Classification
- Optimizing ground state preparation protocols with autoresearch
- Lessons from complex systems science for AI governance
- CayleyPy Growth: Efficient growth computations and hundreds of new conjectures on Cayley graphs (Brief version)
- Reinforced Generation of Combinatorial Structures: Hardness of Approximation
- SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
- Large Language Models as End-to-end Combinatorial Optimization Solvers
- MetaGen: A DSL, Database, and Benchmark for VLM-Assisted Metamaterial Generation
- Improved Constructions and Lower Bounds for Maximally Recoverable Grid Codes
- Large Language Models in Operations Research: Methods, Applications, and Challenges
- Large Language Model Assisted Automated Algorithm Generation and Evolution via Meta-black-box optimization
- Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
- Evolution of Kernels: Automated RISC-V Kernel Optimization with Large Language Models
- EditDuet: A Multi-Agent System for Video Non-Linear Editing
- Autonomous Code Evolution Meets NP-Completeness
- Solve it with EASE
- LLM-Based Instance-Driven Heuristic Bias In the Context of a Biased Random Key Genetic Algorithm
- Artificial intelligence for representing and characterizing quantum systems
- Re-evaluating LLM-based Heuristic Search: A Case Study on the 3D Packing Problem
- Jointly Reinforcing Diversity and Quality in Language Model Generations
- Explicit Constructions of Maximal 3-Zero-Sum-Free Subsets in (ℤ/4ℤ)n
- Scaling Neuro-symbolic Problem Solving: Solver-Free Learning of Constraints and Objectives
- Computer-assisted graph theory: a survey
- Language Models For Generalised PDDL Planning: Synthesising Sound and Programmatic Policies
- ELATE: Evolutionary Language model for Automated Time-series Engineering
- VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization
- HiFo-Prompt: Prompting with Hindsight and Foresight for LLM-based Automatic Heuristic Design
- Discovering Expert-Level Nash Equilibrium Algorithms with Large Language Models
- EvoCut: Strengthening Integer Programs via Evolution-Guided Language Models
- Searching for Privacy Risks in LLM Agents via Simulation
- Route Planning and Online Routing for Quantum Key Distribution Networks
- Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education
- MiGrATe: Mixed-Policy GRPO for Adaptation at Test-Time
- \(X\)-evolve: Solution space evolution powered by large language models
- The Missing Reward: Active Inference in the Era of Experience
- Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies
- Industrial LLM-based Code Optimization under Regulation: A Mixture-of-Agents Approach
- CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search
- How Far Are AI Scientists from Changing the World?
- Automatically discovering heuristics in a complex SAT solver with large language models
- DHEvo: Data-Algorithm Based Heuristic Evolution for Generalizable MILP Solving
- MeLA: A Metacognitive LLM-Driven Architecture for Automatic Heuristic Design
- Pareto-Grid-Guided Large Language Models for Fast and High-Quality Heuristics Design in Multi-Objective Combinatorial Optimization
- Can Language Models Discover Scaling Laws?
- Solving Formal Math Problems by Decomposition and Iterative Reflection
- GENIAL: Generative Design Space Exploration via Network Inversion for Low Power Algorithmic Logic Units
- AI-driven research in pure mathematics and theoretical physics
- Autonomous agentic design for photonics
- AlgoTune: Can Language Models Speed Up General-Purpose Numerical Programs?
- K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model
- How Should We Meta-Learn Reinforcement Learning Algorithms?
- CUDA-L1: Improving CUDA Optimization via Contrastive Reinforcement Learning
- VLMgineer: Vision Language Models as Robotic Toolsmiths
- Large Language Models for Combinatorial Optimization of Design Structure Matrix
- BuildEvo: Designing Building Energy Consumption Forecasting Heuristics via LLM-driven Evolution
- What Remains Human in Mathematics in the Age of AI
- Lean-verified lower bounds for the Shannon capacity of odd cycles
- Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
- Fine-tuning Large Language Model for Automated Algorithm Design
- Self-Improving Language Models for Evolutionary Program Synthesis: A Case Study on ARC-AGI
- Advocate for Complete Benchmarks for Formal Reasoning with Formal/Informal Statements and Formal/Informal Proofs
- Advancing network resilience theories with symbolized reinforcement learning
- Transforming Calabi-Yau Constructions: Generating New Calabi-Yau Manifolds with Transformers
- Large Language Models for Combinatorial Optimization: A Systematic Review
- Behaviour Space Analysis of LLM-driven Meta-heuristic Discovery
- Executable World Models for ARC-AGI-3 in the Era of Coding Agents
- AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench
- Position: A Theory of Deep Learning Must Include Compositional Sparsity
- Performance of LLMs on Stochastic Modeling Operations Research Problems: From Theory to Practice
- The Automated LLM Speedrunning Benchmark: Reproducing NanoGPT Improvements
- Ludax: A GPU-Accelerated Domain Specific Language for Board Games
- EvoVerilog: Large Langugage Model Assisted Evolution of Verilog Code
- GPU Kernel Scientist: An LLM-Driven Framework for Iterative Kernel Optimization
- CoMind: Towards Community-Driven Agents for Machine Learning Engineering
- Programming Quantum Computers with Large Language Models
- Large Language Model-Driven Surrogate-Assisted Evolutionary Algorithm for Expensive Optimization
- HeuriGym: An Agentic Benchmark for LLM-Crafted Heuristics in Combinatorial Optimization
- HeurAgenix: Leveraging LLMs for Solving Complex Combinatorial Optimization Challenges
- LLM Agent for Hyper-Parameter Optimization
- A group-theoretic approach to Shannon capacity of graphs and a limit theorem from lattice packings
- Common Benchmarks Undervalue the Generalization Power of Programmatic Policies
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Automated Heuristic Design for Unit Commitment Using Large Language Models
- Towards Universal Offline Black-Box Optimization via Learning Language Model Embeddings
- Intelligent Design 4.0: Paradigm Evolution Toward the Agentic AI Era
- Can Theoretical Physics Research Benefit from Language Agents?
- AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
- TrajEvo: Designing Trajectory Prediction Heuristics via LLM-driven Evolution
- Optimization Problem Solving Can Transition to Evolutionary Agentic Workflows
- EALG: Evolutionary Adversarial Generation of Language Model-Guided Generators for Combinatorial Optimization
- Improving Generalization of Neural Combinatorial Optimization for Vehicle Routing Problems via Test-Time Projection Learning
- Meta-Optimization and Program Search using Language Models for Task and Motion Planning
- EvoGit: Decentralized Code Evolution via Git-Based Multi-Agent Collaboration
- PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models
- Using Reasoning Models to Generate Search Heuristics that Solve Open Instances of Combinatorial Design Problems
- ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark
- iDSE: Navigating Design Space Exploration in High-Level Synthesis Using LLMs
- LLaMEA-BO: A Large Language Model Evolutionary Algorithm for Automatically Generating Bayesian Optimization Algorithms
- Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives
- Generalizable Heuristic Generation Through LLMs with Meta-Optimization
- RedAHD: Reduction-Based End-to-End Automatic Heuristic Design with Large Language Models
- Learning Extrapolative Sequence Transformations from Markov Chains
- MOOSE-Chem2: Exploring LLM Limits in Fine-Grained Scientific Hypothesis Discovery via Hierarchical Search
- Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators
- Formally Solving Answer-Construction Problems in Lean
- MOOSE-Chem3: Toward Experiment-Guided Hypothesis Ranking via Simulated Experimental Feedback
- DesignX: Human-Competitive Algorithm Designer for Black-Box Optimization
- Dynamic Bundling with Large Language Models for Zero-Shot Inference on Text-Attributed Graphs
- STRCMP: Integrating Graph Structural Priors with Language Models for Combinatorial Optimization
- Long-Horizon Autonomous Architecture Research with a Language-Model Agent: A Behavioural Case Study
- REMS: a unified solution representation, problem modeling and metaheuristic algorithm design for general combinatorial optimization problems
- Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
- Efficient Heuristics Generation for Solving Combinatorial Optimization Problems Using Large Language Models
- Improving Generative Inverse Design of Rectangular Patch Antennas with Test Time Optimization
- CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic Design
- Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models
- From Questions to Clinical Recommendations: Large Language Models Driving Evidence-Based Clinical Decision Making
- Learning Virtual Machine Scheduling in Cloud Computing through Language Agents
- Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch
- CodePDE: An Inference Framework for LLM-driven PDE Solver Generation
- Symbolic Regression with Multimodal Large Language Models and Kolmogorov Arnold Networks
- Visual Evolutionary Optimization on Graph-Structured Combinatorial Problems with MLLMs: A Case Study of Influence Maximization
- Scalable Quantum State Preparation via Large-Language-Model-Driven Discovery
- Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
- HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification
- Improving Coherence and Persistence in Agentic AI for System Optimization
- Towards a new paradigm of scientific discovery with socialized artificial intelligence
- EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
- Do explanations generalize across large reasoning models?
- Semantic Invariance in Agentic AI
- Statistical Proof as a Window into Human-AI Collaboration: Practical Insights and a Community Agenda
- New lower bounds for the degree/diameter problem via interaction with a browser-accessible LLM
- Large Language Model-Driven Full-Component Evolution of Adaptive Large Neighborhood Search
- OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Algorithm Discovery
- VeRO: A Harness for Agents to Optimize Agents
- Act-Observe-Rewrite: Multimodal Coding Agents as In-Context Policy Learners for Robot Manipulation
- BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution
- MLIPilot: LLM-Driven Auto-Research for Machine-Learned Interatomic Potentials
- Self-Improving Language Models with Bidirectional Evolutionary Search
- MadEvolve: Evolutionary Optimization of Trading Systems with Large Language Models
- AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
- Streamlined Constraint Reasoning via CNN Pattern Recognition on Enumerated Solutions
- Automated Kernel Discovery Towards Understanding High-dimensional Bayesian Optimization
- Latent Heuristic Search: Continuous Optimization for Automated Algorithm Design
- DrugSAGE:Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery
- From I/O to Code with Discovery Agent
- Engineering-Oriented Symbolic Regression: LLMs as Physics Agents for Discovery of Simulation-Ready Constitutive Laws
- LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling
- When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning
- SAS-Prompt: Large Language Models as Numerical Optimizers for Robot Self-Improvement
- End-to-end autonomous scientific discovery on a real optical platform
- RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments
- Agentic Symbolic Search: Characterizing PDEs Beyond Hand-crafted Expressions, Meshes, and Neural Networks
- Code evolution for link prediction in complex networks
- Arbor: Tree Search as a Cognition Layer for Autonomous Agents
- Vector Policy Optimization: Training for Diversity Improves Test-Time Search
- Omni-SimpleMem: Autoresearch-Guided Discovery of Lifelong Multimodal Agent Memory
- Can we automatize scientific discovery in the cognitive sciences?
- The Agentic Researcher: A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning
- Prior-Guided Symbolic Regression: Towards Scientific Consistency in Equation Discovery
- Human-AI Co-design for Clinical Prediction Models
- Fitness Landscape of Large Language Model-Assisted Automated Algorithm Search
- AIBuildAI: An AI Agent for Automatically Building AI Models
- AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents
- k-server-bench: Automating Potential Discovery for the k-Server Conjecture
- Meta Context Engineering via Agentic Skill Evolution
- AI for Mathematics: Progress, Challenges, and Prospects
- Beyond Average Performance: Dynamic Instance Clustering and Specialized Algorithm Design in LLM-Assisted Evolutionary Search
- LLM Agents for Combinatorial Efficient Frontiers: Investment Portfolio Optimization
- Vulcan: Instance-specialized, Verifiable Systems Heuristics Through LLM-driven Search
- LoongFlow: Directed Evolutionary Search via a Cognitive Plan-Execute-Summarize Paradigm
- EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement
- Bridging LMS and generative AI: dynamic course content integration (DCCI) for enhancing student satisfaction and engagement via the ask ME assistant
- Frontier AI's Impact on the Cybersecurity Landscape
- Algorithm Discovery With LLMs: Evolutionary Search Meets Reinforcement Learning
- Hierarchical Planning for Complex Tasks with Knowledge Graph-RAG and Symbolic Verification
- CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial Optimization
- Efficient Function Orchestration for Large Language Models
- Evolution of Optimization Algorithms for Global Placement via Large Language Models
- Teaching Humans Subtle Differences with DIFFusion
- Have Large Language Models Learned to Reason? A Characterization via 3-SAT Phase Transition
- The transformative impact of large language models on medical writing and publishing: current applications, challenges and future directions. [europepmc]
- Large language models for human-machine collaborative particle accelerator tuning through natural language. [europepmc]
- Artificial Intelligence Can Drive Sleep Medicine. [europepmc]
- General lightweight framework for vision foundation model supporting multi-task and multi-center medical image analysis. [europepmc]
- RePower: An LLM-driven autonomous platform for power system data-guided research. [europepmc]
- Extending Minds with Generative AI. [europepmc]
- The dual edges of AI: Advancing knowledge while reducing diversity. [europepmc]
- Lessons from complex systems science for AI governance. [europepmc]
- Integrating Large Language Models into Fluid Antenna Systems: A Survey. [europepmc]
- Visual enumeration remains challenging for multimodal generative AI. [europepmc]
- Toward large reasoning models: A survey of reinforced reasoning with large language models. [europepmc]
- Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature? [europepmc]
- Automatically optimizing heuristics for robust scale-free network design via large language models. [europepmc]
- Two-stage prompting framework with predefined verification steps for evaluating diagnostic reasoning tasks on two datasets. [europepmc]
- Streamlining evidence based clinical recommendations with large language models. [europepmc]
- The effect of artificial intelligence-empowered mobile health on psychological distress in women following abortion: protocol for a mixed-methods study. [europepmc]
- Photography Did Not Kill Painting: On Artificial Intelligence and the Future of Academic Medicine. [europepmc]
- An AI system to help scientists write expert-level empirical software. [europepmc]
- AI-discovered tuning laws explain neuronal population code geometry [europepmc]
Related