2006/06/01 by Nicolai Meinshausen, Peter Bühlmann · 2,518 citations
Computer Science · Mathematics · #Bayesian Methods and Mixture Models #Bayesian Modeling and Causal Inference #Conditional independence #Covariance #Covariance matrix #Feature selection #Gaussian #Independence (probability theory) #Lasso (programming language) #Model selection #Multivariate normal distribution #Selection (genetic algorithm) #Statistical Methods and Inference #math.ST #msc:62F12 #msc:62H20 #msc:62J07 #stat.TH
paper · pdf · doi:10.1214/009053606000000281
published in The Annals of Statistics 34(3) (Institute of Mathematical Statistics) · Published at http://dx.doi.org/10.1214/009053606000000281 in the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)
openalex publication_date 2006/06/01 · arxiv created 2006/08/01 · openalex created_date 2016/06/24 · arxiv updated 2016/08/16 · openalex updated_date 2026/08/08
The pattern of zero entries in the inverse covariance matrix of a multivariate normal distribution corresponds to conditional independence restrictions between variables. Covariance selection aims at estimating those structural zeros from data. We show that neighborhood selection with the Lasso is a computationally attractive alternative to standard covariance selection for sparse high-dimensional graphs. Neighborhood selection estimates the conditional independence restrictions separately for each node in the graph and is hence equivalent to variable selection for Gaussian linear models. We show that the proposed neighborhood selection scheme is consistent for sparse high-dimensional graphs. Consistency hinges on the choice of the penalty parameter. The oracle value for optimal prediction does not lead to a consistent neighborhood estimate. Controlling instead the probability of falsely joining some distinct connectivity components of the graph, consistent estimation for sparse graphs is achieved (with exponential rates), even when the number of variables grows as the number of observations raised to an arbitrary power.