2020/04/28 by Aneta Polewko-Klim, Polewko-Klim, Aneta, Witold R. Rudnicki +1
Biochemistry, Genetics and Molecular Biology · #Cancer-related molecular mechanisms research #FOS: Biological sciences #FOS: Computer and information sciences #Gene expression and cancer classification #Genomics (q-bio.GN) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Molecular Biology Techniques and Applications
paper · pdf · doi:10.48550/arxiv.2004.13809
openalex publication_date 2020/04/28 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Discovery of diagnostic and prognostic molecular markers is important and\nactively pursued the research field in cancer research. For complex diseases,\nthis process is often performed using Machine Learning. The current study\ncompares two approaches for the discovery of relevant variables: by application\nof a single feature selection algorithm, versus by an ensemble of diverse\nalgorithms. These approaches are used to identify variables that are relevant\ndiscerning of four cancer types using RNA-seq profiles from the Cancer Genome\nAtlas. The comparison is carried out in two directions: evaluating the\npredictive performance of models and monitoring the stability of selected\nvariables. The most informative features are identified using a four feature\nselection algorithms, namely U-test, ReliefF, and two variants of the MDFS\nalgorithm. Discerning normal and tumor tissues is performed using the Random\nForest algorithm. The highest stability of the feature set was obtained when\nU-test was used. Unfortunately, models built on feature sets obtained from the\nensemble of feature selection algorithms were no better than for models\ndeveloped on feature sets obtained from individual algorithms. On the other\nhand, the feature selectors leading to the best classification results varied\nbetween data sets.\n