2008/06/30 by Hugh A. Chipman, Edward I. George, Robert E. McCulloch · 5 citations
Computer Science · Mathematics · #Bayesian average #Bayesian inference #Bayesian linear regression #Bayesian probability #Boosting (machine learning) #Feature selection #Frequentist inference #Gaussian Processes and Bayesian Inference #Inference #Machine Learning and Algorithms #Markov chain Monte Carlo #Nonparametric regression #Statistical Methods and Inference #stat.AP #stat.ME #stat.ML
paper · pdf · doi:10.1214/09-aoas285
published as Annals of Applied Statistics 2010, Vol. 4, No. 1, 266-298 · Published in at http://dx.doi.org/10.1214/09-AOAS285 the Annals of Applied Statistics (http://www.imstat.org/aoas/) by the Institute of Mathematical Statistics (http://www.imstat.org)
openalex publication_date 2010/03/01 · arxiv created 2010/10/07 · arxiv updated 2010/10/08 · openalex created_date 2016/06/24 · openalex updated_date 2026/08/06
We develop a Bayesian “sum-of-trees” model where each tree is constrained by a regularization prior to be a weak learner, and fitting and inference are accomplished via an iterative Bayesian backfitting MCMC algorithm that generates samples from a posterior. Effectively, BART is a nonparametric Bayesian regression approach which uses dimensionally adaptive random basis elements. Motivated by ensemble methods in general, and boosting algorithms in particular, BART is defined by a statistical model: a prior and a likelihood. This approach enables full posterior inference including point and interval estimates of the unknown regression function as well as the marginal effects of potential predictors. By keeping track of predictor inclusion frequencies, BART can also be used for model-free variable selection. BART’s many features are illustrated with a bake-off against competing methods on 42 different data sets, with a simulation experiment and on a drug discovery classification problem.