2022/10/09 by Cas van der Oord, van der Oord, Cas, Matthias Sachs +7 · 19 citations
Computer Science · Engineering · Materials Science · Mathematics · Physics and Astronomy · #Computational Drug Discovery Methods #Computational Physics (physics.comp-ph) #FOS: Computer and information sciences #FOS: Physical sciences #Fuel Cells and Related Materials #Machine Learning (stat.ML) #Machine Learning in Materials Science #physics.comp-ph #stat.ML
paper · pdf · doi:10.48550/arxiv.2210.04225
21 pages, 11 figures
openalex publication_date 2022/10/09 · arxiv created 2022/11/07 · arxiv updated 2022/11/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Data-driven interatomic potentials have emerged as a powerful class of surrogate models for \it ab initio potential energy surfaces that are able to reliably predict macroscopic properties with experimental accuracy. In generating accurate and transferable potentials the most time-consuming and arguably most important task is generating the training set, which still requires significant expert user input. To accelerate this process, this work presents \it hyperactive learning (HAL), a framework for formulating an accelerated sampling algorithm specifically for the task of training database generation. The key idea is to start from a physically motivated sampler (e.g., molecular dynamics) and add a biasing term that drives the system towards high uncertainty and thus to unseen training configurations. Building on this framework, general protocols for building training databases for alloys and polymers leveraging the HAL framework will be presented. For alloys, ACE potentials for AlSi10 are created by fitting to a minimal HAL-generated database containing 88 configurations (32 atoms each) with fast evaluation times of <100 microsecond/atom/cpu-core. These potentials are demonstrated to predict the melting temperature with excellent accuracy. For polymers, a HAL database is built using ACE, able to determine the density of a long polyethylene glycol (PEG) polymer formed of 200 monomer units with experimental accuracy by only fitting to small isolated PEG polymers with sizes ranging from 2 to 32.