2009/07/03 by Christophe Giraud, Giraud, Christophe, Sylvie Huet +3
Mathematics · #FOS: Mathematics #Statistics Theory (math.ST) #math.ST #stat.TH
paper · pdf · doi:10.48550/arxiv.0907.0619
44 pages
arxiv created 2012/02/15 · arxiv updated 2012/02/17
Applications on inference of biological networks have raised a strong interest in the problem of graph estimation in high-dimensional Gaussian graphical models. To handle this problem, we propose a two-stage procedure which first builds a family of candidate graphs from the data, and then selects one graph among this family according to a dedicated criterion. This estimation procedure is shown to be consistent in a high-dimensional setting, and its risk is controlled by a non-asymptotic oracle-like inequality. The procedure is tested on a real data set concerning gene expression data, and its performances are assessed on the basis of a large numerical study. The procedure is implemented in the R-package GGMselect available on the CRAN.