2019/03/04 by Katharine Turner, Turner, Katharine, Gard Spreemann +1
Computer Science · #Algebraic Topology (math.AT) #Data Management and Algorithms #FOS: Mathematics #Geochemistry and Geologic Mapping #Metric Geometry (math.MG) #Statistics Theory (math.ST) #Topological and Geometric Data Analysis
paper · pdf · doi:10.48550/arxiv.1903.01051
openalex publication_date 2019/03/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Persistent homology allows us to create topological summaries of complex data. In order to analyse these statistically, we need to choose a topological summary and a relevant metric space in which this topological summary exists. While different summaries may contain the same information (as they come from the same persistence module), they can lead to different statistical conclusions since they lie in different metric spaces. The best choice of metric will often be application-specific. In this paper we discuss distance correlation, which is a non-parametric tool for comparing data sets that can lie in completely different metric spaces. In particular we calculate the distance correlation between different choices of topological summaries. We compare some different topological summaries for a variety of random models of underlying data via the distance correlation between the samples. We also give examples of performing distance correlation between topological summaries and other scalar measures of interest, such as a paired random variable or a parameter of the random model used to generate the underlying data. This article is meant to be expository in style, and will include the definitions of standard statistical quantities in order to be accessible to non-statisticians.