Nearly unbiased variable selection under minimax concave penalty
2010/02/19 by Cun-Hui Zhang, Cun‐Hui Zhang · 255 citations
Mathematics · Engineering · #Statistical Methods and Inference #Sparse and Compressive Sensing Techniques #Advanced Statistical Methods and Models
paper · pdf · doi:10.1214/09-aos729
Abstract
We propose MC+, a fast, continuous, nearly unbiased and accurate method of penalized variable selection in high-dimensional linear regression. The LASSO is fast and continuous, but biased. The bias of the LASSO may prevent consistent variable selection. Subset selection is unbiased but computationally costly. The MC+ has two elements: a minimax concave penalty (MCP) and a penalized linear unbiased selection (PLUS) algorithm. The MCP provides the convexity of the penalized loss in sparse regions to the greatest extent given certain thresholds for variable selection and unbiasedness. The PLUS computes multiple exact local minimizers of a possibly nonconvex penalized loss function in a certain main branch of the graph of critical points of the penalized loss. Its output is a continuous piecewise linear path encompassing from the origin for infinite penalty to a least squares solution for zero penalty. We prove that at a universal penalty level, the MC+ has high probability of matching the signs of the unknowns, and thus correct selection, without assuming the strong irrepresentable condition required by the LASSO. This selection consistency applies to the case of p≫n, and is proved to hold for exactly the MC+ solution among possibly many local minimizers. We prove that the MC+ attains certain minimax convergence rates in probability for the estimation of regression coefficients in ℓr balls. We use the SURE method to derive degrees of freedom and Cp-type risk estimates for general penalized LSE, including the LASSO and MC+ estimators, and prove their unbiasedness. Based on the estimated degrees of freedom, we propose an estimator of the noise level for proper choice of the penalty level. For full rank designs and general sub-quadratic penalties, we provide necessary and sufficient conditions for the continuity of the penalized LSE. Simulation results overwhelmingly support our claim of superior variable selection properties and demonstrate the computational efficiency of the proposed method.
Citations
Cited by
- Adaptive deep nonparametric regression from dependent data under covariate shift
- Online Optimization of Difference-of-Convex Compositions with Smooth Mappings
- Frequency Selection in Bayesian Spectral Modeling of Time Series Data with Applications to Wearable Device Measurements
- Evaluating Methods for High-Dimensional Mediation in Metabolomics Data
- Univariate-Guided Sparse Regression
- Kurdyka-Łojasiewicz exponent via inf-projection
- Clustering of longitudinal curves via a penalized method and EM algorithm
- A general framework for deep learning
- On the penalized maximum likelihood estimation of high-dimensional approximate factor model
- Mind the duality gap: safer rules for the Lasso
- Bayesian Feature Extraction using Gaussian and Diffused-gamma Priors for High Dimensional Spatio-Temporal Data
- Difference-of-Convex Elastic Net for Compressed Sensing
- A General Iterative Shrinkage and Thresholding Algorithm for Non-convex Regularized Optimization Problems
- A Practical Guide to Variable Selection in Structural Equation Modeling by Using Regularized Multiple-Indicators, Multiple-Causes Models
- A New Perspective on Debiasing Linear Regressions
- Robust penalized empirical likelihood in high dimensional longitudinal data analysis
- Trace Pursuit: A General Framework for Model-Free Variable Selection
- Sorted Concave Penalized Regression
- A sparse negative binomial mixture model for clustering RNA-seq count\n data
- Alternating Direction Method of Multipliers for A Class of Nonconvex and Nonsmooth Problems with Applications to Background/Foreground Extraction
- Robust Learning for Optimal Treatment Decision with NP-Dimensionality
- Only Train Once: A One-Shot Neural Network Training And Pruning Framework
- Stochastic collocation methods via minimization of Transformed L1 penalty
- Bad estimation, good prediction: the Lasso in dense regimes
- Structured Nonconvex and Nonsmooth Optimization: Algorithms and Iteration Complexity Analysis
- A Likelihood Ratio Framework for High Dimensional Semiparametric Regression
- The dispositional basis of deviance: A multivariate approach to crime and other deviant acts
- Schatten-p Quasi-Norm Regularized Matrix Optimization via Iterative Reweighted Singular Value Minimization
- Renyi Differentially Private ADMM for Non-Smooth Regularized Optimization
- Cross-Semantic Transfer Learning for High-Dimensional Linear Regression
- Shadow splitting methods for nonconvex optimisation: epi-approximation, convergence and saddle point avoidance
- Composite Difference-Max Programs for Modern Statistical Estimation Problems
- Perfect reconstruction of sparse signals using nonconvexity control and one-step RSB message passing
- A preconditioned second-order convex splitting algorithm with extrapolation
- Regularized M-estimators with nonconvexity: Statistical and algorithmic theory for local optima
- On Nonconvex Decentralized Gradient Descent
- Simultaneous Heterogeneity and Reduced-rank Learning for Multivariate Response Regression
- One-Bit Compressed Sensing via One-Shot Hard Thresholding
- A nonmonotone extrapolated proximal gradient-subgradient algorithm beyond global Lipschitz gradient continuity
- Nonconvex sparse regularization for deep neural networks and its optimality
- Numerical optimization for the compatibility constant of the lasso
- Beyond Additivity: Sparse Isotonic Shapley Regression toward Nonlinear Explainability
- Comparing Variable Selection and Model Averaging Methods for Logistic Regression
- Nonconvex Penalized LAD Estimation in Partial Linear Models with DNNs: Asymptotic Analysis and Proximal Algorithms
- Sparse MIMO-OFDM Channel Estimation via RKHS Regularization
- A double iteratively reweighted algorithm for solving group sparse nonconvex optimization models
- Sparse Optimization Problem with s-difference Regularization
- DCA based Algorithm with Extrapolation for Nonconvex Nonsmooth\n Optimization
- Transformed ℓ1 Gradient Regularization for Image Denoising
- Estimation of Spatial and Temporal Autoregressive Effects using LASSO - An Example of Hourly Particulate Matter Concentrations
- On Data Enriched Logistic Regression
- A redescending-weight curvature regularization model for edge-preserving signal denoising
- Strong NP-Hardness for Sparse Optimization with Concave Penalty Functions
- Minimization of the q-ratio sparsity with 1 < q ≤ ∞ for signal recovery
- Optimization landscape of ℓ0-Bregman relaxations
- Multiple-Splitting Projection Test for High-Dimensional Mean Vectors
- Preconditioning to comply with the Irrepresentable Condition
- An efficient proximal algorithm for squared L1 over L2 regularized sparse recovery
- Structured Matrix Scaling for Multi-Class Calibration
- A Stable Lasso
- Efficient Solvers for SLOPE in R, Python, Julia, and C++
- Variable Smoothing Alternating Proximal Gradient Algorithm for Coupled Composite Optimization
- Optimal prediction for sparse linear models? Lower bounds for coordinate-separable M-estimators
- Robust Reduced Rank Regression
- Minimization of Transformed L1 Penalty: Closed Form Representation and Iterative Thresholding Algorithms
- Estimation for bivariate quantile varying coefficient model
- FasTer: Fast Tensor Completion with Nonconvex Regularization
- Doubly Robust Feature Selection with Mean and Variance Outlier Detection and Oracle Properties
- Ideal formulations for constrained convex optimization problems with indicator variables
- ADMM-IDNN: Iteratively Double-reweighted Nuclear Norm Algorithm for\n Group-prior based Nonconvex Compressed Sensing via ADMM
- A New Nonconvex Strategy to Affine Matrix Rank Minimization Problem
- Generalized Linear Model Regression under Distance-to-set Penalties
- Adaptive Threshold Estimation by FDR
- Ranking genetic factors related to age-related maculardegeneration by variable selection confidence sets
- No penalty no tears: Least squares in high-dimensional linear models
- Network Reconstruction via the Minimum Description Length Principle
- Elastic Net Regularization Paths for All Generalized Linear Models
- Provably Convergent Working Set Algorithm for Non-Convex Regularized Regression
- I-LAMM for Sparse Learning: Simultaneous Control of Algorithmic Complexity and Statistical Error
- A ROAD to Classification in High Dimensional Space
- A Simple Correction Procedure for High-Dimensional Generalized Linear\n Models with Measurement Error
- Double Sparsity Kernel Learning with Automatic Variable Selection and Data Extraction
- Improving Network Slimming with Nonconvex Regularization
- Integrative Factor Regression and Its Inference for Multimodal Data Analysis
- Penalized Cox’s proportional hazards model for high-dimensional survival data with grouped predictors
- A constrained L1 minimization approach for estimating multiple Sparse Gaussian or Nonparanormal Graphical Models
- Estimation of oblique structure via penalized likelihood factor analysis
- Gradient descent with nonconvex constraints: local concavity determines convergence
- Stochastic Gradient Descent for Stochastic Doubly-Nonconvex Composite Optimization
- LassoBench: A High-Dimensional Hyperparameter Optimization Benchmark Suite for Lasso
- Effective Proximal Methods for Non-convex Non-smooth Regularized Learning
- The Benefit of Group Sparsity in Group Inference with De-biased Scaled Group Lasso
- Plugging Weight-tying Nonnegative Neural Network into Proximal Splitting Method: Architecture for Guaranteeing Convergence to Optimal Point
- Survival of the fittest Cox model: Pivotal variable selection for time-to-event data
- Improving COVID-19 Forecasting using eXogenous Variables
- A Convexly Constrained LiGME Model and Its Proximal Splitting Algorithm
- Stochastic Difference-of-Convex Optimization with Momentum
- Adaptive Influence Diagnostics in High-Dimensional Regression
- Row-wise Fusion Regularization: An Interpretable Personalized Federated Learning Framework in Large-Scale Scenarios
- Approximate Proximal Operators for Analog Compressed Sensing Using PN-junction Diode
- Tensor Completion via Monotone Inclusion: Generalized Low-Rank Priors Meet Deep Denoisers
- Nonparametric Quantile Regression for Homogeneity Pursuit in Panel Data Models
- Variance Reduced Median-of-Means Estimator for Byzantine-Robust Distributed Inference
- Understanding Notions of Stationarity in Non-Smooth Optimization
- Selective Factor Extraction in High Dimensions
- Statistical Inferences Using Large Estimated Covariances for Panel Data and Factor Models
- Learning Regularization Functionals for Inverse Problems: A Comparative Study
- MSP: A Multi-step Screening Procedure for Sparse Recovery
- Nonconvex and Nonsmooth Sparse Optimization via Adaptively Iterative Reweighted Methods
- Sketching in Bayesian High Dimensional Regression With Big Data Using Gaussian Scale Mixture Priors
- Navigating Sparsities in High-Dimensional Linear Contextual Bandits
- A concave pairwise fusion approach to subgroup analysis
- Sparse-Group Factor Analysis for High-Dimensional Time Series
- Optimality and computational barriers in variable selection under dependence
- Joint Adaptive Penalty for Unbalanced Mediation Pathways
- Variable Selection with Exponential Weights and l0-Penalization
- Sharp RIP bound for sparse signal and low-rank matrix recovery
- Distributed Estimation and Inference with Statistical Guarantees
- SLOPE is Adaptive to Unknown Sparsity and Asymptotically Minimax
- Nested Model Averaging on Solution Path for High-dimensional Linear Regression
- An Efficient ADMM Method for Ratio-Type Nonconvex and Nonsmooth Minimization in Sparse Recovery
- Katalyst: Boosting Convex Katayusha for Non-Convex Problems with a Large\n Condition Number
- Sparse Regularization by Smooth Non-separable Non-convex Penalty Function Based on Ultra-discretization Formula
- Testing Mediation Effects Using Logic of Boolean Matrices
- Group descent algorithms for nonconvex penalized linear and logistic regression models with grouped predictors
- Regularized estimation in sparse high-dimensional time series models
- Model Selection for High Dimensional Quadratic Regression via Regularization
- Penalized integrative analysis under the accelerated failure time model
- APPLE: Approximate Path for Penalized Likelihood Estimators
- Transfer Learning for High-dimensional Linear Regression: Prediction, Estimation, and Minimax Optimality
- Anchored Langevin Algorithms
- Penalty methods for a class of non-Lipschitz optimization problems
- Characterization of Excess Risk for Locally Strongly Convex Population Risk
- Integrative Sparse Partial Least Squares
- Bayesian L(1)/(2) regression
- On the Consistency of the Bias Correction Term of the AIC for the Non-Concave Penalized Likelihood Method
- Independently Interpretable Lasso: A New Regularizer for Sparse Regression with Uncorrelated Variables
- A Rank-Corrected Procedure for Matrix Completion with Fixed Basis Coefficients
- Scalable Bayesian Variable Selection Using Nonlocal Prior Densities in Ultrahigh-Dimensional Settings
- Scalable Hessian-free Proximal Conjugate Gradient Method for Nonconvex and Nonsmooth Optimization
- Nonconvex Regularization for Feature Selection in Reinforcement Learning
- Concave losses for robust dictionary learning
- Matrix Completion with Nonconvex Regularization: Spectral Operators and Scalable Algorithms
- A sparse semismooth Newton based proximal majorization-minimization algorithm for nonconvex square-root-loss regression problems
- Variable Selection for Additive Global Fréchet Regression
- Top-N Recommendation with Novel Rank Approximation
- Distributed Proximal Gradient Algorithm for Partially Asynchronous\n Computer Clusters
- Item Response Theory -- A Statistical Framework for Educational and Psychological Measurement
- Variable Selection Using Relative Importance Rankings
- A √(2)-accelerated FISTA for composite strongly convex problems
- CHIMA: a correlation-aware high-dimensional mediation analysis with its application to the living brain project study
- A Globalized Semismooth Newton Method for Prox-regular Optimization Problems
- A Support Detection and Root Finding Approach for Learning High-dimensional Generalized Linear Models
- Accelerated Block Coordinate Proximal Gradients with Applications in\n High Dimensional Statistics
- Exponential Screening and optimal rates of sparse estimation
- A modified exact penalty approach for general constrained ℓ0-sparse optimization problems
- Gene-Environment Interaction: A Variable Selection Perspective
- Iteratively reweighted ℓ 1 ℓ 1 algorithms with extrapolation
- An Overview on the Estimation of Large Covariance and Precision Matrices
- Sparse PCA with Oracle Property
- Fused Spatial Point Process Intensity Estimation with Varying Coefficients on Complex Constrained Domains
- Estimation and inference for high-dimensional non-sparse models
- Versatile Descent Algorithms for Group Regularization and Variable Selection in Generalized Linear Models
- Strong Rules for Nonconvex Penalties and Their Implications for Efficient Algorithms in High-Dimensional Regression
- Sparse minimum Redundancy Maximum Relevance for feature selection
- Rank-one Convexification for Sparse Regression
- Weakly-Convex Regularization for Magnetic Resonance Image Denoising
- Global Adaptive Generative Adjustment
- Deep Unrolling for Nonconvex Robust Principal Component Analysis
- The restricted consistency property of leave-nv-out cross-validation for high-dimensional variable selection
- Diffusion-Driven High-Dimensional Variable Selection
- An Imputation-Consistency Algorithm for High-Dimensional Missing Data Problems and Beyond
- Adversarial Robust Low Rank Matrix Estimation: Compressed Sensing and Matrix Completion
- Semiparametric Expectile Regression for High-dimensional Heavy-tailed and Heterogeneous Data
- Forward variable selection for sparse ultra-high dimensional varying coefficient models
- HALO: Learning to Prune Neural Networks with Shrinkage
- Inference for High-Dimensional Linear Mixed-Effects Models: A Quasi-Likelihood Approach
- Quantization through Piecewise-Affine Regularization: Optimization and Statistical Guarantees
- Partially linear additive quantile regression in ultra-high dimension
- A Unified Framework for Sparse Relaxed Regularized Regression: SR3
- Algorithmic Bayesian Group Gibbs Selection
- Semiparametric Bayesian Information Criterion for Model Selection in Ultra-high Dimensional Additive Models
- Sparsity Oriented Importance Learning for High-dimensional Linear Regression
- Sparse Regularization: Convergence Of Iterative Jumping Thresholding Algorithm
- Interaction Pursuit with Feature Screening and Selection
- Federated Online Learning for Heterogeneous Multisource Streaming Data
- High-dimensional regression with unknown variance
- Identifiability of the minimum-trace directed acyclic graph and hill climbing algorithms without strict local optima under weakly increasing error variances
- Penalized robust estimators in logistic regression with applications to sparse models
- Multivariate Temporal Point Process Regression
- Generalized Thresholding and Online Sparsity-Aware Learning in a Union\n of Subspaces
- A New Perspective on High Dimensional Confidence Intervals
- On the consistency theory of high dimensional variable screening
- Uncertainty Quantification for Large-Scale Deep Networks via Post-StoNet Modeling
- A Provable Smoothing Approach for High Dimensional Generalized Regression with Applications in Genomics
- Structure Adaptive Elastic-Net
- Penalized Likelihood Estimation in High-Dimensional Time Series Models and its Application
- Ultra High Dimensional Change Point Detection
- On an improvement of LASSO by scaling
- SparseStep: Approximating the Counting Norm for Sparse Regularization
- Variable Selection for Survival Data with A Class of Adaptive Elastic\n Net Techniques
- WEEP: A Differentiable Nonconvex Sparse Regularizer via Weakly-Convex Envelope
- Structural Change in Sparsity
- Expanded Alternating Optimization of Nonconvex Functions with Applications to Matrix Factorization and Penalized Regression
- Lasso Penalization for High-Dimensional Beta Regression Models: Computation, Analysis, and Inference
- Fast Low-Rank Matrix Learning with Nonconvex Regularization
- Run-and-Inspect Method for Nonconvex Optimization and Global Optimality Bounds for R-Local Minimizers
- Innovated scalable efficient estimation in ultra-large Gaussian graphical models
- Learning Latent Features with Pairwise Penalties in Low-Rank Matrix Completion
- A New Integrative Learning Framework for Integrating Multiple Secondary Outcomes into Primary Outcome Analysis: A Case Study on Liver Health
- DuRIN: A Deep-unfolded Sparse Seismic Reflectivity Inversion Network
- Sparsity-Agnostic Lasso Bandit
- Weak Signal Identification and Inference in Penalized Model Selection
- A Variational Approach on Level sets and Linear Convergence of Variable Bregman Proximal Gradient Method for Nonconvex Optimization Problems
- Statistical Inference in High-dimensional Generalized Linear Models with Streaming Data
- A note relating ridge regression and OLS p-values to preconditioned sparse penalized regression
- A proximal dual semismooth Newton method for computing zero-norm penalized QR estimator
- Large-Scale Low-Rank Matrix Learning with Nonconvex Regularizers
- Stability selection enables robust learning of partial differential equations from limited noisy data
- Tensor Generalized Estimating Equations for Longitudinal Imaging Analysis
- Parallel subgroup analysis of high-dimensional data via M-regression
- Estimation And Selection Via Absolute Penalized Convex Minimization And Its Multistage Adaptive Applications
- Partial Penalized Likelihood Ratio Test under Sparse Case
- Least-Square Approximation for a Distributed System
- Scalable Algorithms for the Sparse Ridge Regression
- Extended Comparisons of Best Subset Selection, Forward Stepwise Selection, and the Lasso
- Low Rank Regularization: A Review
- Lasso Meets Horseshoe : A Survey
- A General Theory of Hypothesis Tests and Confidence Regions for Sparse High Dimensional Models
- M-estimation with the Trimmed l1 Penalty
- Penalized Variable Selection for Multi-center Competing Risks Data
- Between hard and soft thresholding: optimal iterative thresholding algorithms
- Path Following and Empirical Bayes Model Selection for Sparse Regression
- Variable Selection with Second-Generation P-Values
- A Two-Stage Penalized Least Squares Method for Constructing Large\n Systems of Structural Equations
- Simple structure estimation via prenet penalization
- Algorithmic Versatility of SPF-regularization Methods
- Nearly optimal Bayesian Shrinkage for High Dimensional Regression
- Spectral Analysis of High-dimensional Time Series
- Sparse Solution of Underdetermined Linear Equations via Adaptively Iterative Thresholding
- Unified Scalable Equivalent Formulations for Schatten Quasi-Norms
- Confidence Intervals for Low-Dimensional Parameters in High-Dimensional Linear Models
- Accelerated Stochastic Algorithms for Nonconvex Finite-sum and Multi-block Optimization
- Distribution Regression
- Distributed statistical optimization for non-randomly stored big data with application to penalized learning
- Capped Lp approximations for the composite L0 regularization problem
- Comparisons of penalized least squares methods by simulations
- Big Data Analysis Using Shrinkage Strategies
- Sparse Fisher's Linear Discriminant Analysis for Partially Labeled Data
- Sparse recovery via nonconvex regularized M-estimators over ℓq-balls
- Optimal Statistical Inference for Individualized Treatment Effects in High-dimensional Models
- Inference for high-dimensional instrumental variables regression
- Pathwise Coordinate Optimization for Sparse Learning: Algorithm and Theory
- A Novel Approach for Fast Detection of Multiple Change Points in Linear Models
- Relaxed Sparse Eigenvalue Conditions for Sparse Estimation via Non-convex Regularized Regression
Related