2013/07/29 by Stan Hatko, Hatko, Stan
Computer Science · Mathematics · #Artificial intelligence #Bayesian Methods and Mixture Models #Computer science #Curse of dimensionality #Dimensionality reduction #FOS: Computer and information sciences #Machine Learning (stat.ML) #Machine learning #Mathematics #Medical Image Segmentation Techniques #Reduction (mathematics) #Statistical Methods and Inference #stat.ML
paper · pdf · doi:10.48550/arxiv.1307.8333
Winter 2013 Honours research project done under the supervision of Dr. Vladimir Pestov at the University of Ottawa; 45 pages
arxiv created 2013/07/29 · openalex publication_date 2013/07/29 · arxiv updated 2013/08/01 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28
In this project we further investigate the idea of reducing the dimensionality of datasets using a Borel isomorphism with the purpose of subsequently applying supervised learning algorithms, as originally suggested by my supervisor V. Pestov (in 2011 Dagstuhl preprint). Any consistent learning algorithm, for example kNN, retains universal consistency after a Borel isomorphism is applied. A series of concrete examples of Borel isomorphisms that reduce the number of dimensions in a dataset is provided, based on multiplying the data by orthogonal matrices before the dimensionality reducing Borel isomorphism is applied. We test the accuracy of the resulting classifier in a lower dimensional space with various data sets. Working with a phoneme voice recognition dataset, of dimension 256 with 5 classes (phonemes), we show that a Borel isomorphic reduction to dimension 16 leads to a minimal drop in accuracy. In conclusion, we discuss further prospects of the method.