2018/06/21 by Hsiang Hsu, Hsu, Hsiang, Salman Salamatian +3
Agricultural and Biological Sciences · Biochemistry, Genetics and Molecular Biology · Physics and Astronomy · #Complex Network Analysis Techniques #FOS: Computer and information sciences #Gene expression and cancer classification #Information Theory (cs.IT) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Sensory Analysis and Statistical Methods
paper · pdf · doi:10.48550/arxiv.1806.08449
openalex publication_date 2018/06/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Correspondence analysis (CA) is a multivariate statistical tool used to\nvisualize and interpret data dependencies by finding maximally correlated\nembeddings of pairs of random variables. CA has found applications in fields\nranging from epidemiology to social sciences; however, current methods do not\nscale to large, high-dimensional datasets. In this paper, we provide a novel\ninterpretation of CA in terms of an information-theoretic quantity called the\nprincipal inertia components. We show that estimating the principal inertia\ncomponents, which consists in solving a functional optimization problem over\nthe space of finite variance functions of two random variable, is equivalent to\nperforming CA. We then leverage this insight to design novel algorithms to\nperform CA at an unprecedented scale. Particularly, we demonstrate how the\nprincipal inertia components can be reliably approximated from data using deep\nneural networks. Finally, we show how these maximally correlated embeddings of\npairs of random variables in CA further play a central role in several learning\nproblems including visualization of classification boundary and training\nprocess, and underlying recent multi-view and multi-modal learning methods.\n