2016/02/16 by Nathaniel E. Helwig, Ping Ma, Helwig, Nathaniel E. +1 · 1 citation
Computer Science · Mathematics · #Computation (stat.CO) #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Neural Networks and Applications #Statistical Methods and Inference
paper · pdf · doi:10.48550/arxiv.1602.05208
openalex publication_date 2016/02/16 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28
In the current era of big data, researchers routinely collect and analyze data of super-large sample sizes. Data-oriented statistical methods have been developed to extract information from super-large data. Smoothing spline ANOVA (SSANOVA) is a promising approach for extracting information from noisy data; however, the heavy computational cost of SSANOVA hinders its wide application. In this paper, we propose a new algorithm for fitting SSANOVA models to super-large sample data. In this algorithm, we introduce rounding parameters to make the computation scalable. To demonstrate the benefits of the rounding parameters, we present a simulation study and a real data example using electroencephalography data. Our results reveal that (using the rounding parameters) a researcher can fit nonparametric regression models to very large samples within a few seconds using a standard laptop or tablet computer.