vix.ing · top · new · best · stats · spec

Projected support points: a new method for high-dimensional data\n reduction

2017/08/23 by Simon Mak, V. Roshan Joseph, Mak, Simon +1 · 1 citation
Computer Science · Engineering · Mathematics · Medicine · #3D Shape Modeling and Analysis #Advanced Numerical Analysis Techniques #Algorithm #Artificial intelligence #Bayesian probability #Big data #Computation #Computer science #Data mining #Data point #Data set #Dimensionality reduction #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Kernel (algebra) #Machine learning #Markov chain Monte Carlo #Mathematical Approximation and Integration #Mathematics #Medical Imaging Techniques and Applications #Methodology (stat.ME) #Reduction (mathematics) #stat.ME

paper · pdf · doi:10.48550/arxiv.1708.06897

openalex publication_date 2017/08/23 · arxiv created 2018/06/02 · arxiv updated 2018/06/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

In an era where big and high-dimensional data is readily available, data\nscientists are inevitably faced with the challenge of reducing this data for\nexpensive downstream computation or analysis. To this end, we present here a\nnew method for reducing high-dimensional big data into a representative point\nset, called projected support points (PSPs). A key ingredient in our method is\nthe so-called sparsity-inducing (SpIn) kernel, which encourages the\npreservation of low-dimensional features when reducing high-dimensional data.\nWe begin by introducing a unifying theoretical framework for data reduction,\nconnecting PSPs with fundamental sampling principles from experimental design\nand Quasi-Monte Carlo. Through this framework, we then derive sparsity\nconditions under which the curse-of-dimensionality in data reduction can be\nlifted for our method. Next, we propose two algorithms for one-shot and\nsequential reduction via PSPs, both of which exploit big data subsampling and\nmajorization-minimization for efficient optimization. Finally, we demonstrate\nthe practical usefulness of PSPs in two real-world applications, the first for\ndata reduction in kernel learning, and the second for reducing Markov Chain\nMonte Carlo (MCMC) chains.\n

Citations

Cited by

Related