2015/11/12 by Émilie Devijver, Emilie Devijver, Mélina Gallopin +2 · 2 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Algorithm #Bioinformatics and Genomic Networks #Block matrix #Cluster analysis #Computer science #Covariance #Covariance matrix #Diagonal #Estimation of covariance matrices #FOS: Computer and information sciences #FOS: Mathematics #Gaussian #Gene Regulatory Network Analysis #Gene expression and cancer classification #Graphical model #Lasso (programming language) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematical optimization #Mathematics #Matrix (chemical analysis) #Minimax #Model selection #Sample size determination #Statistics #Statistics Theory (math.ST) #Stochastic block model #cs.LG #math.ST #stat.ML #stat.TH
paper · pdf · doi:10.48550/arxiv.1511.04033
published in arXiv (Cornell University) (Cornell University) · Accepted in JASA
openalex publication_date 2015/11/12 · arxiv created 2016/09/29 · arxiv updated 2016/09/30 · openalex created_date 2022/10/04 · openalex updated_date 2026/07/28
Gaussian graphical models are widely utilized to infer and visualize networks of dependencies between continuous variables. However, inferring the graph is difficult when the sample size is small compared to the number of variables. To reduce the number of parameters to estimate in the model, we propose a non-asymptotic model selection procedure supported by strong theoretical guarantees based on an oracle inequality and a minimax lower bound. The covariance matrix of the model is approximated by a block-diagonal matrix. The structure of this matrix is detected by thresholding the sample covariance matrix, where the threshold is selected using the slope heuristic. Based on the block-diagonal structure of the covariance matrix, the estimation problem is divided into several independent problems: subsequently, the network of dependencies between variables is inferred using the graphical lasso algorithm in each block. The performance of the procedure is illustrated on simulated data. An application to a real gene expression dataset with a limited sample size is also presented: the dimension reduction allows attention to be objectively focused on interactions among smaller subsets of genes, leading to a more parsimonious and interpretable modular network.