2023/07/11 by Tianyu Wang, Wang, Tianyu, Jiashuo Liu +5 · 1 voice · 15 citations
Computer Science · Decision Sciences · Social Sciences · #Artificial Intelligence (cs.AI) #Big Data Technologies and Applications #FOS: Computer and information sciences #Insurance, Mortality, Demography, Risk Management #Machine Learning (cs.LG) #cs.AI #cs.LG #demographic modeling and climate adaptation
paper · pdf · doi:10.48550/arxiv.2307.05284
openalex publication_date 2023/07/11 · arxiv published 2023/07/11 · openalex created_date 2023/07/13 · arxiv updated 2026/06/03 · openalex updated_date 2026/07/28
Different distribution shifts require different interventions, and algorithms must be grounded in the specific shifts they address. However, methodological development for robust algorithms typically relies on structural assumptions that lack empirical validation. Advocating for an empirically grounded data-driven approach to algorithm development, we build an empirical testbed comprising natural shifts across 8 tabular datasets, 172 distribution pairs over 45 methods and 90,000 method configurations encompassing empirical risk minimization and distributionally robust optimization (DRO) methods. We find Y|X-shifts are most prevalent in our testbed, in stark contrast to the heavy focus on X (covariate)-shifts in the ML literature, and that the performance of robust algorithms is no better than that of vanilla methods. To understand why, we conduct an in-depth empirical analysis of DRO methods and find that underlooked implementation details -- such as the choice of underlying model class (e.g., LightGBM) and hyperparameter selection -- have a bigger impact on performance than the ambiguity set or its radius. We illustrate via case studies how a data-driven, inductive understanding of distribution shifts can provide a new approach to algorithm development.