2020/12/31 by V. Roshan Joseph, Akhil Vakayil · 198 citations
Computer Science · Mathematics · #Categorical variable #Face and Expression Recognition #Imbalanced Data Classification Techniques #Machine Learning and Data Classification #Point (geometry) #Regression #Regression analysis #Training set #cs.LG #k-nearest neighbors algorithm #stat.ML
paper · pdf · doi:10.1080/00401706.2021.1921037
published in Technometrics 64(2), 166-176 (Taylor & Francis)
openalex created_date 2021/01/05 · arxiv created 2021/03/19 · openalex publication_date 2021/04/28 · arxiv updated 2021/05/10 · openalex updated_date 2026/08/06
In this article, we propose an optimal method referred to as SPlit for splitting a dataset into training and testing sets. SPlit is based on the method of support points (SP), which was initially developed for finding the optimal representative points of a continuous distribution. We adapt SP for subsampling from a dataset using a sequential nearest neighbor algorithm. We also extend SP to deal with categorical variables so that SPlit can be applied to both regression and classification problems. The implementation of SPlit on real datasets shows substantial improvement in the worst-case testing performance for several modeling methods compared to the commonly used random splitting procedure.