vix.ing · top · new · best · stats

Automated design of collective variables using supervised machine learning

2018/09/07 by Mohammad M. Sultan, Vijay S. Pande · 145 citations
Computer Science · #Advanced Multi-Objective Optimization Algorithms #Artificial intelligence #Computer science #Evolutionary Algorithms and Applications #Machine learning #Software Engineering Research

paper · doi:10.1063/1.5029972

published in The Journal of Chemical Physics 149(9), 094106 (American Institute of Physics)

openalex publication_date 2018/09/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01

Abstract

Selection of appropriate collective variables (CVs) for enhancing sampling of molecular simulations remains an unsolved problem in computational modeling. In particular, picking initial CVs is particularly challenging in higher dimensions. Which atomic coordinates or transforms there of from a list of thousands should one pick for enhanced sampling runs? How does a modeler even begin to pick starting coordinates for investigation? This remains true even in the case of simple two state systems and only increases in difficulty for multi-state systems. In this work, we solve the “initial” CV problem using a data-driven approach inspired by the field of supervised machine learning (SML). In particular, we show how the decision functions in SML algorithms can be used as initial CVs (SMLcv) for accelerated sampling. Using solvated alanine dipeptide and Chignolin mini-protein as our test cases, we illustrate how the distance to the support vector machines’ decision hyperplane, the output probability estimates from logistic regression, the outputs from shallow or deep neural network classifiers, and other classifiers may be used to reversibly sample slow structural transitions. We discuss the utility of other SML algorithms that might be useful for identifying CVs for accelerating molecular simulations.

Citations

Cited by

Related