2026/02/25 by Andrea Apicella, Francesco Isgrò, Francesco Isgro +2 · 1 voice
Computer Science · #Adversarial Robustness in Machine Learning #Cross-validation #Domain Adaptation and Few-Shot Learning #Early stopping #Generalization #Model selection #Model validation #Selection (genetic algorithm) #Set (abstract data type) #Stochastic Gradient Optimization Techniques #Test set #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2602.22107
openalex publication_date 2026/02/25 · arxiv published 2026/02/25 · arxiv updated 2026/02/25 · openalex created_date 2026/02/27 · openalex updated_date 2026/07/28
Despite the extensive literature on training loss functions, the evaluation of generalization on the validation set remains underexplored. In this work, we conduct a systematic empirical and statistical study of how the validation criterion used for model selection affects test performance in neural classifiers, with attention to early stopping. Using fully connected networks on standard benchmarks under k-fold evaluation, we compare: (i) early stopping with patience and (ii) post-hoc selection over all epochs (i.e. no early stopping). Models are trained with cross-entropy, C-Loss, or PolyLoss; the model parameter selection on the validation set is made using accuracy or one of the three loss functions, each considered independently. Three main findings emerge. (1) Early stopping based on validation accuracy performs worst, consistently selecting checkpoints with lower test accuracy than both loss-based early stopping and post-hoc selection. (2) Loss-based validation criteria yield comparable and more stable test accuracy. (3) Across datasets and folds, any single validation rule often underperforms the test-optimal checkpoint. Overall, the selected model typically achieves test-set performance statistically lower than the best performance across all epochs, regardless of the validation criterion. Our results suggest avoiding validation accuracy (in particular with early stopping) for parameter selection, favoring loss-based validation criteria.