vix.ing · top · new · best · stats

Robust high dimensional factor models with applications to statistical machine learning

2018/08/12 by Jianqing Fan, Fan, Jianqing, Kaizheng Wang +5 · 12 citations
Computer Science · Engineering · Mathematics · #Artificial intelligence #Blind Source Separation Techniques #Computer science #Curse of dimensionality #Data mining #FOS: Computer and information sciences #FOS: Mathematics #Factor (programming language) #Factor analysis #Inference #Machine Learning (stat.ML) #Machine learning #Mathematics #Methodology (stat.ME) #Principal component analysis #Random projection #Sparse and Compressive Sensing Techniques #Statistical Methods and Inference #Statistical hypothesis testing #Statistical inference #Statistical learning theory #Statistical model #Statistics #Statistics Theory (math.ST) #Support vector machine #math.ST #stat.ME #stat.ML #stat.TH

paper · pdf · doi:10.48550/arxiv.1808.03889

published in arXiv (Cornell University) (Cornell University) · 41 pages

arxiv created 2018/08/12 · openalex publication_date 2018/08/12 · arxiv updated 2018/08/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06

Abstract

Factor models are a class of powerful statistical models that have been widely used to deal with dependent measurements that arise frequently from various applications from genomics and neuroscience to economics and finance. As data are collected at an ever-growing scale, statistical machine learning faces some new challenges: high dimensionality, strong dependence among observed variables, heavy-tailed variables and heterogeneity. High-dimensional robust factor analysis serves as a powerful toolkit to conquer these challenges. This paper gives a selective overview on recent advance on high-dimensional factor models and their applications to statistics including Factor-Adjusted Robust Model selection (FarmSelect) and Factor-Adjusted Robust Multiple testing (FarmTest). We show that classical methods, especially principal component analysis (PCA), can be tailored to many new problems and provide powerful tools for statistical estimation and inference. We highlight PCA and its connections to matrix perturbation theory, robust statistics, random projection, false discovery rate, etc., and illustrate through several applications how insights from these fields yield solutions to modern challenges. We also present far-reaching connections between factor models and popular statistical learning problems, including network analysis and low-rank matrix recovery.

Citations

Cited by

Related