2018/02/28 by Mohammad M. Sultan, Vijay S. Pande, Sultan, Mohammad M. +1 · 3 citations
Materials Science · Biochemistry, Genetics and Molecular Biology · Chemistry · #Machine Learning in Materials Science #Protein Structure and Dynamics #Mass Spectrometry Techniques and Applications
paper · pdf · doi:10.48550/arxiv.1802.10510
Selection of appropriate collective variables for enhancing sampling of\nmolecular simulations remains an unsolved problem in computational biophysics.\nIn particular, picking initial collective variables (CVs) is particularly\nchallenging in higher dimensions. Which atomic coordinates or transforms there\nof from a list of thousands should one pick for enhanced sampling runs? How\ndoes a modeler even begin to pick starting coordinates for investigation? This\nremains true even in the case of simple two state systems and only increases in\ndifficulty for multi-state systems. In this work, we solve the initial CV\nproblem using a data-driven approach inspired by the filed of supervised\nmachine learning. In particular, we show how the decision functions in\nsupervised machine learning (SML) algorithms can be used as initial CVs\n(SMLcv) for accelerated sampling. Using solvated alanine dipeptide and\nChignolin mini-protein as our test cases, we illustrate how the distance to the\nSupport Vector Machines' decision hyperplane, the output probability estimates\nfrom Logistic Regression, the outputs from deep neural network classifiers, and\nother classifiers may be used to reversibly sample slow structural transitions.\nWe discuss the utility of other SML algorithms that might be useful for\nidentifying CVs for accelerating molecular simulations.\n