2020/06/05 by Andreas Vitalis, Vitalis, Andreas
Computer Science · Economics, Econometrics and Finance · #68R10 (Primary) 68W10 (Secondary) #Complex Systems and Time Series Analysis #Data Structures and Algorithms (cs.DS) #Data Visualization and Analytics #Distributed #E.1 #FOS: Computer and information sciences #G.2.2 #I.5.3 #Parallel #Time Series Analysis and Forecasting #and Cluster Computing (cs.DC)
paper · pdf · doi:10.48550/arxiv.2006.04940
openalex publication_date 2020/06/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Today, very large amounts of data are produced and stored in all branches of\nsociety including science. Mining these data meaningfully has become a\nconsiderable challenge and is of the broadest possible interest. The size, both\nin numbers of observations and dimensionality thereof, requires data mining\nalgorithms to possess time complexities with both variables that are linear or\nnearly linear. One such algorithm, see Comput. Phys. Commun. 184, 2446-2453\n(2013), arranges observations into a sequence called the progress index. The\nprogress index steps through distinct regions of high sampling density\nsequentially. By means of suitable annotations, it allows a compact\nrepresentation of the behavior of complex systems, which is encoded in the\noriginal data set. The only essential parameter is a notion of distance between\nobservations. Here, we present the shared memory parallelization of the key\nstep in constructing the progress index, which is the calculation of an\napproximation of the minimum spanning tree of the complete graph of\nobservations. We demonstrate that excellent parallel efficiencies are obtained\nfor up to 72 logical (CPU) cores. In addition, we introduce three conceptual\nadvances to the algorithm that improve its controllability and the\ninterpretability of the progress index itself.\n