vix.ing · top · new · best · stats · spec

Combining predictions from linear models when training and test inputs\n differ

2014/06/24 by Thijs van Ommen, van Ommen, Thijs
Computer Science · Mathematics · #62F07 (Primary) 62C12 #62J05 (Secondary) #Advanced Statistical Methods and Models #Artificial intelligence #Bayesian probability #Computer science #Covariate #Divergence (linguistics) #Econometrics #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Linear model #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine learning #Mathematics #Methodology (stat.ME) #Model selection #Selection (genetic algorithm) #Set (abstract data type) #Statistical Methods and Inference #Test (biology) #Test set #Training set #cs.LG #msc:62C12 #msc:62F07 #msc:62J05 #stat.ME #stat.ML

paper · pdf · doi:10.48550/arxiv.1406.6200

12 pages, 2 figures. To appear in Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence (UAI2014). This version includes the supplementary material (regularity assumptions, proofs)

arxiv created 2014/06/24 · openalex publication_date 2014/06/24 · arxiv updated 2014/06/25 · openalex created_date 2025/10/24 · openalex updated_date 2026/07/28

Abstract

Methods for combining predictions from different models in a supervised\nlearning setting must somehow estimate/predict the quality of a model's\npredictions at unknown future inputs. Many of these methods (often implicitly)\nmake the assumption that the test inputs are identical to the training inputs,\nwhich is seldom reasonable. By failing to take into account that prediction\nwill generally be harder for test inputs that did not occur in the training\nset, this leads to the selection of too complex models. Based on a novel,\nunbiased expression for KL divergence, we propose XAIC and its special case\nFAIC as versions of AIC intended for prediction that use different degrees of\nknowledge of the test inputs. Both methods substantially differ from and may\noutperform all the known versions of AIC even when the training and test inputs\nare iid, and are especially useful for deterministic inputs and under covariate\nshift. Our experiments on linear models suggest that if the test and training\ninputs differ substantially, then XAIC and FAIC predictively outperform AIC,\nBIC and several other methods including Bayesian model averaging.\n

Citations

Related