2009/11/01 by Jason L. Loeppky, Jerome Sacks, William J. Welch · 760 citations
Computer Science · Decision Sciences · Mathematics · #Advanced Multi-Objective Optimization Algorithms #Algorithm #Code (set theory) #Computer experiment #Computer science #Dimension (graph theory) #Gaussian #Gaussian Processes and Bayesian Inference #Key (lock) #Machine learning #Mathematics #Optimal Experimental Design Methods #Process (computing) #Programming language #Sample (material) #Sample size determination #Sensitivity (control systems) #Simulation #Source code #Statistics #Theoretical computer science #Variable (mathematics) #Variables
paper · doi:10.1198/tech.2009.08040
published in Technometrics 51(4), 366-376 (Taylor & Francis)
openalex publication_date 2009/11/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
We provide reasons and evidence supporting the informal rule that the number of runs for an effective initial computer experiment should be about 10 times the input dimension. Our arguments quantify two key characteristics of computer codes that affect the sample size required for a desired level of accuracy when approximating the code via a Gaussian process (GP). The first characteristic is the total sensitivity of a code output variable to all input variables; the second corresponds to the way this total sensitivity is distributed across the input variables, specifically the possible presence of a few prominent input factors and many impotent ones (i.e., effect sparsity). Both measures relate directly to the correlation structure in the GP approximation of the code. In this way, the article moves toward a more formal treatment of sample size for a computer experiment. The evidence supporting these arguments stems primarily from a simulation study and via specific codes modeling climate and ligand activation of G-protein.