vix.ing · top · new · best · stats · spec

Gravity-Bench-v1: A Benchmark on Gravitational Physics Discovery for Agents

2025/01/30 by Koblischke, Nolan, Jang, Hyunseok, Kristen Menou +3 · 3 citations
Computer Science · #Algorithms and Data Compression #Artificial Intelligence (cs.AI) #Computability, Logic, AI Algorithms #Computational Physics (physics.comp-ph) #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #FOS: Physical sciences #Instrumentation and Methods for Astrophysics (astro-ph.IM)

paper · pdf · doi:10.48550/arxiv.2501.18411

openalex publication_date 2025/01/30 · openalex created_date 2025/02/01 · openalex updated_date 2026/07/28

Abstract

Modern science emerged from reasoning over repeatedly-observed planetary motions. We present Gravity-Bench-v1, an environment-based benchmark that challenges AI agents on tasks that parallel this historical development. Gravity-Bench-v1 evaluates agents on the discovery of physics concealed within a dynamic environment, using rigorous gravitational dynamics simulations. Gravity-Bench includes out-of-distribution cases, i.e. with physics that deviates from the real world, to evaluate true scientific generalization capabilities. Agents must plan to collect data within an experimental budget and must perform a dynamic form of data analysis and reasoning to solve tasks efficiently. Our benchmark admits an open-ended space of solutions. Reference solutions for each task are provided to calibrate AI performance against human expertise. Technically at an upper-undergraduate level, our benchmark proves challenging to baseline AI agents. Gravity-Bench-v1 and planned extensions should help map out AI progress towards scientific discovery capabilities.

Cited by

Related