2020/08/17 by Javad Rahimipour Anaraki, Anaraki, Javad Rahimipour, Saeed Samet +1
Biochemistry, Genetics and Molecular Biology · Computer Science · #68P27 #Cryptography and Security (cs.CR) #Data Mining Algorithms and Applications #FOS: Computer and information sciences #Gene expression and cancer classification #I.5.1 #I.5.2 #Machine Learning (cs.LG) #Privacy-Preserving Technologies in Data #Rough Sets and Fuzzy Logic
paper · pdf · doi:10.48550/arxiv.2008.07664
openalex publication_date 2020/08/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Feature selection is the process of sieving features, in which informative\nfeatures are separated from the redundant and irrelevant ones. This process\nplays an important role in machine learning, data mining and bioinformatics.\nHowever, traditional feature selection methods are only capable of processing\ncentralized datasets and are not able to satisfy today's distributed data\nprocessing needs. These needs require a new category of data processing\nalgorithms called privacy-preserving feature selection, which protects users'\ndata by not revealing any part of the data neither in the intermediate\nprocessing nor in the final results. This is vital for the datasets which\ncontain individuals' data, such as medical datasets. Therefore, it is rational\nto either modify the existing algorithms or propose new ones to not only\nintroduce the capability of being applied to distributed datasets, but also act\nresponsibly in handling users' data by protecting their privacy. In this paper,\nwe will review three privacy-preserving feature selection methods and provide\nsuggestions to improve their performance when any gap is identified. We will\nalso propose a privacy-preserving feature selection method based on the rough\nset feature selection. The proposed method is capable of processing both\nhorizontally and vertically partitioned datasets in two- and multi-parties\nscenarios.\n