2025/12/01 by Ahmad Ali, Ali, Ahmad
Biochemistry, Genetics and Molecular Biology · Computer Science · Materials Science · #Categorical variable #Computational Drug Discovery Methods #Entropy (arrow of time) #Histogram #Machine Learning in Materials Science #Protein Structure and Dynamics #Regression #Regression analysis #Representation (politics) #Scalar (mathematics)
paper · pdf · doi:10.48550/arxiv.2512.01160
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/12/01 · openalex created_date 2025/12/03 · openalex updated_date 2026/07/28
Density Functional Theory (DFT) is a widely used computational method for estimating the energy and behavior of molecules. Machine Learning Interatomic Potentials (MLIPs) are models trained to approximate DFT-level energies and forces at dramatically lower computational cost. Many modern MLIPs rely on a scalar regression formulation; given information about a molecule, they predict a single energy value and corresponding forces while minimizing absolute error with DFT's calculations. In this work, we explore a multi-class classification formulation that predicts a categorical distribution over energy/force values, providing richer supervision through multiple targets. Most importantly, this approach offers a principled way to quantify model uncertainty. In particular, our method predicts a histogram of the energy/force distribution, converts scalar targets into histograms, and trains the model using cross-entropy loss. Our results demonstrate that this categorical formulation can achieve absolute error performance comparable to regression baselines. Furthermore, this representation enables the quantification of epistemic uncertainty through the entropy of the predicted distribution, offering a measure of model confidence absent in scalar regression approaches.