2026/06/16 by Arkaprava Banerjee, Kunal Roy · 1 voice
Computer Science · #Machine Learning and Data Classification #Explainable Artificial Intelligence (XAI) #Computational Physics and Python Applications
paper · pdf · doi:10.26434/chemrxiv.15004777/v1
openalex publication_date 2026/06/16 · openalex created_date 2026/06/17 · openalex updated_date 2026/07/15
Selecting a single best machine learning regression model from a set of competing models can be challenging. While models selected based on cross-validation performance do not guarantee good predictions on external data, models selected solely on external validation performance do not ascertain precise predictions for other external sets. Therefore, we propose three quantitative metrics to guide modelers in selecting the best model. Three quantitative datasets of varying sizes and complexities were considered. Each dataset was randomly split into a modeling set and an independent external test set. The modeling set was further split thrice to generate training and validation sets. Various machine learning models were developed and validated against the validation sets. Our proposed metrics were computed for each model using only training and validation set performances. The novel metric values guided the selection of the best models, which also demonstrated expected performance on the independent external test set.