vix.ing · top · new · best · stats

A framework for a generalisation analysis of machine-learned interatomic potentials

2022/09/12 by Christoph Ortner, Yangshuai Wang, Ortner, Christoph +1 · 1 citation
Biochemistry, Genetics and Molecular Biology · Computer Science · Engineering · Materials Science · Mathematics · #Advanced Electron Microscopy Techniques and Applications #Algorithm #Artificial intelligence #Computer science #Cover (algebra) #Current (fluid) #Electron and X-Ray Spectroscopy Techniques #Engineering #Geometry #Interatomic potential #Machine Learning in Materials Science #Mathematics #Mechanical engineering #Molecular dynamics #Physics #Point (geometry) #Quantum mechanics #Section (typography) #Space (punctuation) #Statistical physics #Training set #cs.NA #math.NA

paper · pdf · doi:10.48550/arxiv.2209.05366

published in arXiv (Cornell University) (Cornell University)

arxiv created 2022/09/12 · openalex publication_date 2022/09/12 · arxiv updated 2022/09/13 · openalex created_date 2022/10/01 · openalex updated_date 2026/08/05

Abstract

Machine-learned interatomic potentials (MLIPs) and force fields (i.e. interaction laws for atoms and molecules) are typically trained on limited data-sets that cover only a very small section of the full space of possible input structures. MLIPs are nevertheless capable of making accurate predictions of forces and energies in simulations involving (seemingly) much more complex structures. In this article we propose a framework within which this kind of generalisation can be rigorously understood. As a prototypical example, we apply the framework to the case of simulating point defects in a crystalline solid. Here, we demonstrate how the accuracy of the simulation depends explicitly on the size of the training structures, on the kind of observations (e.g., energies, forces, force constants, virials) to which the model has been fitted, and on the fit accuracy. The new theoretical insights we gain partially justify current best practices in the MLIP literature and in addition suggest a new approach to the collection of training data and the design of loss functions.

Cited by

Related