2023/05/07 by Huang, Dongming, Tian, Songtao, Lin, Qian
#62C20 (Secondary) #62J02 (Primary) 62H12 #FOS: Mathematics #Statistics Theory (math.ST)
paper · doi:10.48550/arxiv.2305.04340
In this work, we address the longstanding puzzle that Sliced Inverse Regression (SIR) often performs poorly for sufficient dimension reduction when the structural dimension d (the dimension of the central space) exceeds 4. We first show that in the multiple index model Y=f( P \boldsymbolX)+ε where \boldsymbolX is a p-standard normal vector, ε is an independent noise, and P is a projection operator from \mathbb Rp to \mathbb Rd, if the link function f follows the law of a Gaussian process, then with high probability, the d-th eigenvalue λd of Cov[𝔼(\boldsymbolX| Y)] satisfies λd≤ C e-θd for some positive constants C and θ. We then focus on the low signal regime where λd can be arbitrarily small and not larger than d-8.1, and prove that the minimax risk of estimating the central space is lower bounded by \fracdpnλd. Combining these two results, we provide a convincing explanation for the poor performance of SIR when d is large, a phenomenon that has perplexed researchers for nearly three decades. The technical tools developed here may be of independent interest for studying other sufficient dimension reduction methods.